Reconsidering Novelty in Bayesian Confirmation — Epoche C1
The question, stated concretely On the night of 23 September 1846 Johann Galle turned a telescope to the place in the sky where Urbain Le Verrier's calculations said an unknown planet must be, and found Neptune within about a degree of the predicted position. The question this essay examines is whether the fact that the observation came after the prediction, rather than before it, adds anything to the support the observation gave Le Verrier's hypothesis — and if so, what a probabilistic theory of confirmation has to say about it. The intuition that it does add something is widely shared, and the compressed version of this argument proposed that Bayesian confirmation theory should be supplemented with a distinct 'surprise' element to capture it. The conclusion reached here is different and, I think, better for the intuition: no supplement is needed, because surprise is already in the formalism, exactly and derivably, once one stops looking for it in the wrong quantity. What genuinely resists formal treatment is something else, and identifying it correctly is more useful than adding a term. The machinery, and what cancels in it Some notation first, since everything turns on it. Let $H$ be a hypothesis, $E$ a proposition reporting evidence, and $K$ everything else one takes for granted. $P(H \mid K)$ is the prior : the probability of $H$ before $E$ is learnt. $P(H \mid E \wedge K)$ is the posterior : the probability after. $P(E \mid H \wedge K)$ is the likelihood of the evidence on the hypothesis — note that this is the probability of the evidence given the hypothesis, not the reverse, and confusing the two directions is the commonest error in this area. Bayes' theorem relates them: $$P(H \mid E \wedge K) \;=\; \frac{P(E \mid H \wedge K)\,P(H \mid K)}{P(E \mid K)}.$$ Bayesian confirmation is then defined as probability-raising: $E$ confirms $H$ relative to $K$ when $P(H \mid E \wedge K)$ exceeds $P(H\mid K)$. Now write the same theorem for $\neg H$, the denial of the hypothesis, and divide one equation by the other. The term $P(E\mid K)$ appears in both denominators and cancels, leaving $$\frac{P(H \mid E \wedge K)}{P(\neg H \mid E \wedge K)} \;=\; \frac{P(E \mid H \wedge K)}{P(E \mid \neg H \wedge K)} \times \frac{P(H \mid K)}{P(\neg H \mid K)}.$$ In words: posterior odds equal the Bayes factor — the ratio of the two likelihoods, written $\Lambda$ below — multiplied by prior odds. This is the form worth carrying, because it isolates the contribution of the evidence in a single quantity, and because of what is missing from it. The marginal probability of the evidence, $P(E\mid K)$, does not appear. Whatever role surprise plays in confirmation, it cannot be played by the unconditional improbability of the evidence, because that quantity has cancelled out. What a low probability of the evidence does and does not buy This is the first place where the standard story needs correcting, and the correction matters because the usual reply to the novelty intuition — that surprise is just a psychological shadow cast by a low $P(E)$ — is not quite right. By the law of total probability, the marginal probability of the evidence decomposes into contributions from the two branches: $$P(E\mid K) \;=\; P(E\mid H\wedge K)\,P(H\mid K) \;+\; P(E\mid \neg H\wedge K)\,P(\neg H\mid K).$$ A low value of $P(E\mid K)$ is therefore consistent with two quite different situations. It can arise because $P(E\mid\neg H\wedge K)$ is minute while $P(E\mid H \wedge K)$ is large, in which case $\Lambda$ is enormous and the evidence is powerfully confirming. Or it can arise because $E$ is improbable on every hypothesis under consideration, in which case both likelihoods are small, $\Lambda$ is near one, and the evidence confirms nothing at all despite being maximally surprising. An unheralded meteor strike is improbable; it does not confirm anything about planetary dynamics. The compressed argument noticed something real here — its example of a bizarre but previously considered outcome that fails to impress — and diagnosed it as a defect in the formalism. It is not. If an outcome was previously considered, some rival hypothesis already predicted it, so $P(E\mid\neg H\wedge K)$ is not small, so $\Lambda$ is not large. The intuition is vindicated by the likelihood ratio rather than requiring a supplement to it. It is worth seeing how large $\Lambda$ was in the Neptune case, since the number can be derived rather than gestured at. Take $H$ to be Le Verrier's hypothesis — a planet of the calculated mass at the calculated longitude — and $E$ the observation of a previously uncatalogued planetary body within one degree of that longitude. On $H$, the likelihood is close to $1$: given the planet is there and bright enough, a competent observer with the ephemeris finds it. On $\neg H$, model the alternative crudely: if no such planet exists, then any object that might be mistaken for one lies with roughly uniform probability somewhere along the band of sky in which planets are found, a strip of about four degrees' width running the full $360$ degrees around the ecliptic, giving about $1440$ square degrees. A circle of angular radius one degree covers $\pi \approx 3.1$ square degrees. So $P(E\mid \neg H \wedge K) \approx 3.1/1440 \approx 2\times 10^{-3}$, and $\Lambda \approx 500$. The band rather than the whole sky is the right reference class because a new planet would necessarily lie near the ecliptic, and $\neg H$ must be granted that much too; using the whole sky, $41\,253$ square degrees, would inflate $\Lambda$ by a factor of about thirty and take credit for information Le Verrier did not supply. A Bayes factor of a few hundred is enough to move a hypothesis from serious speculation to near certainty, which is what happened. Surprise, made exact The proposal that confirmation theory needs a distinct measure of surprise can now be assessed, and the assessment is a two-line derivation. Return to Bayes' theorem and rearrange: $$\fra