Bayesian Epistemology and the Empirical Testing of Scientific Hypotheses — Epoche C1
A failed prediction does not say what failed In 1859 the French astronomer Urbain Le Verrier reported that the perihelion of Mercury — the point of its orbit closest to the Sun — advances about 43 arcseconds per century more than Newtonian gravitation allows. That residual is worth pausing over, because it shows how such a figure is arrived at. Measured against the fixed stars the perihelion appears to move by roughly 5,600 arcseconds per century; of that, about 5,025 arcseconds is not a motion of Mercury at all but of the coordinate frame, the slow precession of the Earth's equinoxes; a further 532 arcseconds or so is the gravitational tug of the other planets, chiefly Venus, Jupiter and the Earth. Subtract both and 43 arcseconds per century remain. Something was wrong. Nothing in the observation said what. This is the difficulty the philosophy of science calls the Duhem problem, and its logic is elementary. A theory $T$ never yields an observable prediction by itself. It does so only in company with auxiliary assumptions $A$ — that the instrument is calibrated, that no unmodelled body is present, that the sample is representative. What is testable is the conjunction. So if $T \wedge A$ entails a prediction $P$ and we observe $\neg P$, valid inference delivers $$\neg (T \wedge A), \quad \text{that is} \quad \neg T \vee \neg A,$$ and a disjunction is not a verdict. Modus tollens — the rule that from 'if $X$ then $Y$' and 'not $Y$' one may infer 'not $X$' — refutes the package and remains silent about the parts. Le Verrier's own career illustrates both outcomes. Confronted in the 1840s with anomalies in the orbit of Uranus, he blamed the auxiliary assumption that the known planets were all the planets, computed where an unknown body would have to be, and Johann Galle found Neptune in 1846 within a degree of the predicted position. Confronted with Mercury, he made the same move, postulated an intra-Mercurial planet named Vulcan — and no such planet exists. The residual was resolved only in 1915, when Einstein's general theory of relativity yielded a perihelion advance for Mercury of the observed size from a revision of $T$ rather than of $A$. The same inference pattern, the same distinguished astronomer, opposite correct answers. Pierre Duhem drew the general moral in The Aim and Structure of Physical Theory (French original 1906), arguing among other things that there are no crucial experiments: the Foucault measurement of 1850, which found light slower in water than in air and was taken to decide between the wave and the corpuscular theories, decides nothing on its own, since the corpuscular theory could in principle be repaired. One point of scope deserves stating plainly, because the standard label 'the Quine–Duhem thesis' obscures it. Duhem confined his claim to theoretical physics and expressly denied that it applied everywhere in science. It was W. V. O. Quine, in 'Two Dogmas of Empiricism' (1951), who extended holism to the whole of knowledge, holding that any statement can be held true come what may if we make drastic enough adjustments elsewhere in the system, up to and including revision of logic. The two theses have different strengths and different consequences, and the argument below engages Duhem's, which is the one that concerns experimental practice. Degrees of belief, and the theorem that governs them The Bayesian response begins by refusing to work with acceptance and rejection at all, and replacing them with graded confidence. A degree of belief is not a private feeling but a quantity defined by what one is prepared to do: on Frank Ramsey's analysis in 'Truth and Probability' (1931), your degree of belief in a proposition is the price at which you would be willing to buy or sell a bet that pays a unit if it is true. Ramsey showed that anyone whose betting prices violate the axioms of probability can be offered a set of bets each of which they accept and which together guarantee them a loss. Coherence, not introspective accuracy, is what forces degrees of belief into the probability calculus. Given that, the updating rule is Bayes's theorem, which is a rearrangement of the definition of conditional probability and so is not in dispute: $$P(H \mid E) = \frac{P(E \mid H)\, P(H)}{P(E)}, \qquad P(E) = P(E \mid H)\,P(H) + P(E \mid \neg H)\,P(\neg H).$$ Here $P(H)$ is the prior, one's confidence in the hypothesis before the observation; $P(E \mid H)$ is the likelihood, the probability the hypothesis assigns to the observation; $P(H \mid E)$ is the posterior, the confidence after it; and $P(E)$ is the total probability of the observation, computed by averaging its probability over the hypothesis and its negation. The second equation is the one that matters for the Duhem problem, because it is where the alternatives to $H$ enter the calculation. Blame is allocated, not assigned The insight, developed by Jon Dorling in 'Bayesian Personalism, the Methodology of Scientific Research Programmes, and Duhem's Problem' (1979), is that the disjunction $\neg T \vee \neg A$ is a dead end only for deductive logic. Probabilistically, an anomaly divides its impact between the conjuncts, and how it divides depends on quantities that scientists routinely have opinions about. To see this we need four cases rather than two, since $T$ and $A$ can each be true or false. Take the essay's ecological setting and make it definite. A team studies regeneration in selectively logged rainforest. $H$ is the substantive hypothesis that stem density recovers to a stated fraction of unlogged density within fifteen years of a logging ban; $A$ is the auxiliary that the plot protocol measures stem density without systematic bias. Together they predict that the surveyed plots will exceed a threshold density; the evidence $E$ is that the plots fall well short of it. The inputs are these, and each is stated with its reason. $P(H) = 0.7$: the hypothesis rests on published chronosequences from comparable forest, but n