The Rate Function, Not the Variance, Governs Rare Events — Epoche C2
Ask how likely it is that the average of a thousand independent measurements lands far from its expected value, and the reply will usually invoke the central limit theorem. The reasoning is that the theorem describes the distribution of the average, so one need only read off a Gaussian tail. This is the misconception the present essay is concerned with. The central limit theorem is a statement about a shrinking window around the mean; it is silent about the region where rare events live. The statement that fills the gap is Cramér's theorem, and the object it produces — the rate function — turns out to contain the variance as a single special case. What the central limit theorem actually controls Let $X_1, X_2, \dots$ be independent random variables with the same distribution, with mean $\mu$ and finite variance $\sigma^2$, and write $S_n = X_1 + \dots + X_n$ for their sum. The central limit theorem says that the standardised sum $(S_n - n\mu)/(\sigma\sqrt{n})$ converges in distribution to a standard normal variable. Convergence in distribution means that the probability of any fixed interval converges. The interval is fixed in the standardised variable, which is the whole difficulty: a fixed interval there corresponds to a window around $\mu$ of width proportional to $n^{-1/2}$, and that window closes as $n$ grows. Now consider the event $S_n/n \ge \mu + \delta$ for a fixed $\delta > 0$ — an average that misses by a fixed margin. In standardised units this is a departure of $\delta\sqrt{n}/\sigma$ standard deviations, which diverges. The theorem is being asked about a region that is escaping to infinity in exactly the coordinate in which the theorem is stated. It therefore says nothing. The point can be made sharper with the Berry–Esseen theorem, which bounds how far the exact distribution function $F_n$ of the standardised sum lies from the normal one $\Phi$: with $\rho = \mathbb{E}|X_1 - \mu|^3$ the third absolute central moment, $\sup_x |F_n(x) - \Phi(x)| \le C\rho/(\sigma^3\sqrt{n})$ for an absolute constant $C$. The bound is absolute , not relative. If the quantity being estimated is of size $10^{-50}$ and the guaranteed accuracy is $10^{-2}$, the guarantee is worthless — the two numbers differ by forty-eight orders of magnitude. This is why the misconception is invisible in practice: nothing in the theorem is false, it is simply not being asked a question it answers. Cramér's theorem supplies the missing statement Define the cumulant generating function $\Lambda(\lambda) = \log \mathbb{E}[e^{\lambda X_1}]$, and assume it is finite for all $\lambda$ in some neighbourhood of zero — Cramér's condition. Define the rate function as its Legendre–Fenchel transform, $$I(x) = \sup_{\lambda \in \mathbb{R}} \big( \lambda x - \Lambda(\lambda) \big),$$ and Cramér's theorem states that for $x > \mu$, $$\lim_{n \to \infty} \frac{1}{n} \log \mathbb{P}\!\left( \frac{S_n}{n} \ge x \right) = -I(x).$$ Half of this is one line. For $\lambda > 0$, Markov's inequality applied to $e^{\lambda S_n}$ gives $\mathbb{P}(S_n \ge nx) \le e^{-n\lambda x}\,\mathbb{E}[e^{\lambda S_n}] = e^{-n(\lambda x - \Lambda(\lambda))}$, because the summands are independent and the expectation factorises. Optimising over $\lambda$ produces $e^{-nI(x)}$; the supremum in the definition of $I$ is simply the best Chernoff bound. The matching lower bound is obtained by tilting the measure — reweighting the distribution by $e^{\lambda^* X}$ so that $x$ becomes the new mean, after which the event is typical and the law of large numbers applies to the tilted variables. Three properties follow immediately. First, $\Lambda$ is convex, so $I$ is convex as a supremum of affine functions of $x$. Second, $\Lambda(0) = 0$ and $\Lambda'(0) = \mu$, so the supremum at $x = \mu$ is attained at $\lambda^* = 0$ and $I(\mu) = 0$. Third, $I(x) > 0$ elsewhere. The law of large numbers is now a corollary rather than a separate theorem. Since $I$ is positive away from $\mu$, the probability that $|S_n/n - \mu| \ge \varepsilon$ is bounded by $2e^{-n\eta}$ with $\eta = \min\{I(\mu+\varepsilon), I(\mu-\varepsilon)\} > 0$. That sequence is summable in $n$, so by the Borel–Cantelli lemma only finitely many such deviations occur, which is the strong law. Where the Gaussian answer is right, and where it is not For normal summands the two pictures agree, and this is why the misconception survives. If $X_1$ is normal with mean $\mu$ and variance $\sigma^2$ then $\Lambda(\lambda) = \mu\lambda + \sigma^2\lambda^2/2$; setting the derivative of $\lambda x - \Lambda(\lambda)$ to zero gives $\lambda^* = (x-\mu)/\sigma^2$ and hence $I(x) = (x-\mu)^2/(2\sigma^2)$. The naive Gaussian extrapolation is exact. Take instead a fair coin, so $\mu = 1/2$ and $\sigma^2 = 1/4$. Here $\Lambda(\lambda) = \log\big((1+e^{\lambda})/2\big)$, and the transform yields $I(x) = x\log x + (1-x)\log(1-x) + \log 2$, which is the Kullback–Leibler divergence of a coin with bias $x$ from the fair coin — the form Sanov's theorem predicts for empirical frequencies. At $x = 0.75$, in natural logarithms, $I = 0.75\log 0.75 + 0.25\log 0.25 + \log 2 = -0.2158 - 0.3466 + 0.6931 = 0.1308$ nats. The Gaussian guess is $(0.25)^2/(2 \times 0.25) = 0.125$. The gap is only $0.0058$ per unit of $n$, but it is multiplied by $n$: at $n = 1000$ the exponents are $-130.8$ and $-125.0$, so the Gaussian answer is too large by $e^{5.81} \approx 3.3 \times 10^{2}$ — about two and a half orders of magnitude, comparing $10^{2.52}$ with $10^{0}$. Two caveats are worth stating plainly. Cramér's theorem fixes only the exponential scale; the polynomial prefactor requires a saddle-point refinement of the tilted measure, which is the content of the Bahadur–Rao expansion. And if Cramér's condition fails — if $\mathbb{E}[e^{\lambda X_1}] = \infty$ for every $\lambda > 0$, as for a Pareto tail — then $I$ vanishes on the whole half-line above the mean and the exponential scale is simply the wrong scale. For such subexponential su