Predictive Coding Forbids Outcomes; the Free-Energy Principle Does Not — Epoche C2
Predictive coding asserts that the cortex contains two interleaved populations of cells — one carrying predictions sent downwards, one carrying the mismatch between prediction and input sent upwards — with the gain on the mismatch cells set by an estimate of how reliable that mismatch is. The free-energy principle asserts that any system which persists in a bounded set of states will appear to minimise a bound on the improbability of its sensory states. The two are almost always introduced in the same breath, and the pairing is presented as a success story: one principle explains what the brain is for, and neurophysiology confirms it. This essay argues that the pairing is the source of the confusion rather than the achievement. The two claims differ in what they assert, in what would count against them, and therefore in how they should be judged. Taking them apart is not a concession to either side; it is the only way to say what the evidence bears on. The quantity both proposals start from Both begin with the same expression, so the divergence has to be located after it rather than in it. Let $o$ be the sensory data an organism receives and $\vartheta$ the hidden causes it must infer. Exact inference requires the posterior $p(\vartheta \mid o)$, which for any realistic generative model is intractable because the normalising integral over $\vartheta$ cannot be evaluated. Variational inference replaces it with an approximating density $q(\vartheta)$ drawn from a tractable family and defines the variational free energy $$F[q, o] = \mathbb{E}_{q}\!\left[\ln q(\vartheta) - \ln p(o, \vartheta)\right] = -\ln p(o) + D_{\mathrm{KL}}\!\left(q(\vartheta)\,\|\,p(\vartheta \mid o)\right),$$ where $D_{\mathrm{KL}}$ is the Kullback–Leibler divergence. The second equality is one substitution — write $p(o,\vartheta) = p(\vartheta \mid o)\,p(o)$ and take the constant $\ln p(o)$ outside the expectation — and since $D_{\mathrm{KL}} \ge 0$ it shows $F$ to be an upper bound on the surprisal $-\ln p(o)$, tight exactly when $q$ equals the posterior. A second regrouping, this time splitting $p(o,\vartheta) = p(o \mid \vartheta)\,p(\vartheta)$, gives $$F[q, o] = D_{\mathrm{KL}}\!\left(q(\vartheta)\,\|\,p(\vartheta)\right) - \mathbb{E}_{q}\!\left[\ln p(o \mid \vartheta)\right],$$ the difference between how far the inferred density has been dragged from the prior and how well it explains the data. Minimising $F$ over $q$ therefore tightens an approximation while penalising departures from prior expectation; minimising it over actions that change $o$ reduces surprisal itself. Nothing so far is a claim about brains. It is the standard variational apparatus, and it is available to anyone. From a bound to a claim about cells Predictive coding is what one obtains by adding two assumptions that are not forced by the mathematics above, and this is the step at which an empirical hypothesis enters. First, approximate $q$ by a Gaussian centred on its mode, which is the Laplace approximation; its effect is that the whole density is carried by one number per hidden cause, so a population of cells can plausibly represent it. Second, assume the generative model is hierarchical, each level predicting the level below through a nonlinearity, with additive Gaussian noise at every level. Under those two assumptions $F$ reduces to a sum of squared prediction errors weighted by their precisions, $$F \simeq \tfrac{1}{2}\sum_{i}\left(\Pi_{i}\,\varepsilon_{i}^{2} - \ln \Pi_{i}\right) + \mathrm{const},$$ where $\varepsilon_i = \mu_i - g_i(\mu_{i+1})$ is the mismatch at level $i$ between what the level above predicts through the function $g_i$ and what the level actually represents, and $\Pi_i$ is the precision, the inverse variance, of that mismatch. Every term earns its place. The squared error is the exponent of the Gaussian; the $-\ln \Pi_i$ is its normalising constant, and it is not decoration. Without it, the expression would be minimised by driving every precision to zero, which is to say by ignoring all sensory evidence. With it, differentiating with respect to $\Pi_i$ gives $\tfrac{1}{2}\left(\varepsilon_i^2 - \Pi_i^{-1}\right) = 0$, so the optimal precision is $\Pi_i = 1/\varepsilon_i^2$: the gain a level assigns to its own error signal is the inverse of the size that error has been running at. This is the formal core of what the literature calls precision weighting, and it is why the same term is invoked in accounts of attention and of gain control. Gradient descent on the same expression gives the update rule. Since $\mu_i$ appears in the error at its own level and, through $g_{i-1}$, in the error at the level below, $$\dot{\mu}_i \;\propto\; -\frac{\partial F}{\partial \mu_i} \;=\; \left(\frac{\partial g_{i-1}}{\partial \mu_i}\right)^{\!\top}\Pi_{i-1}\varepsilon_{i-1} \;-\; \Pi_i \varepsilon_i,$$ which is exactly the Rao–Ballard scheme: each representational unit is driven upward by the precision-weighted error beneath it and corrected by the error at its own level. What Rao and Ballard themselves reported in 1999 was not this derivation but a working model. They trained a hierarchical network of this form on natural image patches and found that units in it reproduced end-stopping — the reduction in a neuron's response when an oriented bar is extended beyond its classical receptive field — as a consequence of the extension being predictable from the surround, not as a built-in property. That is a specific, checkable claim about why a specific extra-classical effect occurs. Once stated in these terms the theory is a claim about cells. It says that some neurons carry predictions, others carry errors, that the two are anatomically distinguishable, and that the gain on error units tracks an estimate of precision. A further and logically separable proposal, set out by Bastos and colleagues in 2012, assigns the two populations to different cortical laminae: superficial pyramidal cells carrying feedforward error and expressing gamma-