Inductive Probability in Scientific and Forensic Contexts — Epoche C1
The rodeo, and why a probability above one half will not convict A thousand people are seated in a rodeo arena and the organiser has sold four hundred and ninety-nine tickets. Nobody was counted at the gate, and no other evidence exists. Pick any spectator at random and sue him for the price of admission. The probability that he did not pay is $501/1000 = 0.501$, which clears the civil standard of proof in English law — the balance of probabilities, meaning more likely than not. Yet no court would give judgment against him, and the intuition that it should not is very firm. This is the paradox of the gatecrasher, and L. Jonathan Cohen used it in The Probable and the Provable of 1977 to open a case against the application of ordinary probability to legal and forensic proof. He put forward an alternative he called inductive or Baconian probability, and the original version of this essay argued that Cohen's alternative handles the practical difficulties of forensic and epidemiological reasoning better than the Bayesian account does. That claim can be assessed, but only after Cohen's system has been stated as a system, which the original did not do, and after the Bayesian side has been given the resources it actually possesses, which the original also did not do. Two of the original's specific charges turn out to be misdirected, and one of Cohen's arguments turns out to be considerably stronger than the essay made it look. Bayes' theorem, and the odds form that forensic practice actually uses Bayes' theorem follows in two lines from the definition of conditional probability. Since $P(H \wedge E)$ can be written either as $P(H \mid E)P(E)$ or as $P(E \mid H)P(H)$, equating the two and dividing gives $$ P(H \mid E) = \frac{P(E \mid H)\,P(H)}{P(E)} . $$ Here $H$ is a hypothesis, $E$ the evidence, $P(H)$ the prior probability of the hypothesis before the evidence is considered, $P(E \mid H)$ the likelihood — the probability of getting that evidence if the hypothesis is true — and $P(H \mid E)$ the posterior, the revised degree of belief. The theorem is not in dispute; it is a consequence of the axioms. The original essay's principal practical complaint was that Bayesian inference forces an arbitrary prior probability of guilt onto a suspect. That complaint is answered within the Bayesian framework, and the answer is what forensic science already does. Write $H_p$ for the prosecution's hypothesis and $H_d$ for the defence's, apply the theorem to each, and divide one by the other. The term $P(E)$ appears in both denominators and cancels, leaving $$ \frac{P(H_p \mid E)}{P(H_d \mid E)} \;=\; \frac{P(E \mid H_p)}{P(E \mid H_d)} \times \frac{P(H_p)}{P(H_d)} . $$ In words: the posterior odds equal the likelihood ratio multiplied by the prior odds. The middle term — the ratio of how expected the evidence is under each hypothesis — is the only one that depends on the scientific findings, and it can be computed from population data without any view about guilt. The Association of Forensic Science Providers made this division of labour the professional standard in its 2009 statement on evaluative opinion: the expert reports the likelihood ratio and states explicitly that the prior odds are a matter for the court. The prior is therefore not forced onto anyone; it is deliberately left where the Bayesian analysis says it belongs. This does not dissolve the gatecrasher paradox, because in that case the likelihood ratio is all the evidence there is. It does mean that the original essay's central practical argument for preferring Cohen is directed at a version of Bayesianism that forensic practice abandoned. Base-rate neglect is a fact about people, not a defect in the theory The original essay's second charge was that base-rate neglect is a bias "that Bayesian frameworks struggle to fully accommodate normatively". This inverts the relationship, and the inversion is worth correcting because it affects what the evidence shows. Base-rate neglect is the tendency to underweight general population frequencies in favour of specific case information. Kahneman and Tversky demonstrated it in 1973: subjects were told that a personality description had been drawn from a pool of either seventy engineers and thirty lawyers or thirty engineers and seventy lawyers, and their judgements about the described individual's profession were nearly the same in both conditions. The base rate, which differed by a factor of more than five in the odds, made almost no difference to what they said. Notice what makes that a finding at all. It is a finding because the base rate should have mattered, and the standard by which it should have mattered is Bayes' theorem. Base-rate neglect is defined by reference to the Bayesian norm; the norm is the instrument of measurement, not the thing found wanting. A descriptive failure of human reasoners is not a normative failure of the standard they fail to meet. The magnitude of the error is worth having in front of one. Take a screening test for a condition present in one per cent of the population, with a sensitivity of ninety per cent — the probability of a positive result given the condition — and a false-positive rate of nine per cent. Then $$ P(\text{condition} \mid +) = \frac{0.90 \times 0.01}{0.90 \times 0.01 + 0.09 \times 0.99} = \frac{0.0090}{0.0981} \approx 0.09 . $$ Fewer than one positive result in ten is a true positive, a figure that physicians asked this question routinely put an order of magnitude too high. Gigerenzer and Hoffrage established in 1995 what the mistake actually consists in, and their result tells against the original essay's use of it. They presented the same information in natural frequencies rather than probabilities: of a thousand people, ten have the condition, of whom nine test positive; of the remaining nine hundred and ninety, about eighty-nine test positive; so nine of ninety-eight positives are genuine. On that presentation the proportion of respondents reaching