Causal Understanding and Steiner's Criterion for Explanatory Proof — Epoche C2
The analogy this essay defends, and a correction to its source Two epidemiologists can both assert that smoking causes lung cancer, both be right, both be justified, and yet differ in a way that the propositional content of their assertion does not record: one can say what a confounding factor would have to look like to account for the observed association without smoke doing any causal work, and the other cannot. This essay argues that the difference is the whole of what we should mean by causal understanding, and that the resources for saying so precisely already exist in the philosophy of mathematics, in an account of what makes one proof of a theorem explanatory when another proof of the same theorem is not. That account is Mark Steiner's, and the earlier version of this essay mis-cited it twice. It named the author 'Michael Steiner' and gave the source as Mathematical Knowledge (1978). The author is Mark Steiner; Mathematical Knowledge is his 1975 book, which is chiefly a defence of mathematical realism on broadly Quinean grounds; and the account of explanatory proof appears three years later in a separate paper, 'Mathematical Explanation' (1978). The correction is not pedantry. The 1978 paper does not merely observe that mathematicians distinguish proofs that convince from proofs that explain — everyone grants that — it proposes a criterion for the distinction, and the criterion is what makes the transposition to causal inference more than a metaphor. Steiner's criterion, stated exactly A proof, on any account, establishes that a theorem is true. Mathematicians nonetheless say routinely of one valid proof that it explains the theorem and of another equally valid proof that it does not, and the problem Steiner sets himself is to say what property the first has. He rejects several natural candidates. Explanatoriness is not greater generality, since a proof can be more general and less illuminating. It is not visualisability, since many explanatory proofs in analysis and number theory are unvisualisable. It is not proceeding directly from definitions, since a proof can unwind definitions mechanically and explain nothing. His positive proposal has two clauses. First, an explanatory proof makes reference to a characterising property of an entity or structure named in the theorem — a property that singles that entity out within some family of related entities — and the proof must visibly turn on that property, so that one can see where the result comes from. Second, and this is the operational clause, the proof must deform : substituting a different member of the family, and with it a different characterising property, must convert the proof into a proof of the corresponding theorem about that member. A proof that survives deformation into an array of related theorems has exhibited the dependence of the result on the property; a proof that shatters when you vary the setup has not shown you what the result depends on, however conclusively it has shown you that the result holds. The standard illustration is the sum of the first $n$ positive integers. Proved by induction, one verifies the base case and then verifies that if $1 + 2 + \cdots + k = k(k+1)/2$ then adding $k+1$ gives $(k+1)(k+2)/2$; the computation is correct and the theorem is established. Nothing in it depends on a property peculiar to the sequence $1, 2, 3, \ldots$, and nothing in it deforms: told the answer for a different sequence, the induction will verify that too, but it will not have found it. The pairing argument is different. Write the sum forwards and backwards, add the two rows term by term, and every column is the same total $n+1$, because the sequence has a constant difference and so the increase down one row exactly cancels the decrease up the other. There are $n$ columns, so $2S = n(n+1)$ and $S = n(n+1)/2$ — for $n = 100$, $100 \times 101 / 2 = 5050$. The characterising property here is the constant difference, and the proof turns on it explicitly, which is what lets the proof deform. Replace the sequence by any arithmetic progression $a, a+d, \ldots, a+(n-1)d$ and the same pairing gives every column the total $2a + (n-1)d$, whence $$S = \frac{n\,[\,2a + (n-1)d\,]}{2}.$$ Setting $a = 1$ and $d = 1$ returns $n(n+1)/2$; setting $a = 1$ and $d = 2$ gives $n(2 + 2n - 2)/2 = n^2$, the sum of the first $n$ odd numbers, which for $n = 4$ reads $1 + 3 + 5 + 7 = 16$. One proof, correctly deformed, has produced a family of theorems, and in doing so has displayed exactly which feature of the original sequence the original result rested on. That display is what Steiner means by explanation. Why this bears on understanding at all The philosophically load-bearing consequence is easy to state and easy to miss. Explanatoriness, on Steiner's criterion, is a property of a derivation and not of the proposition derived. The two proofs above establish a numerically identical theorem; whatever distinguishes them is therefore invisible at the level of propositional content. It follows that if understanding tracks explanation at all, understanding cannot be individuated by the propositions the understander believes, however true and however justified. Something about the route must enter. This is the point at which the earlier version of this essay was right but unsupported, and it is worth marking that the same conclusion has been argued directly in epistemology rather than by analogy. Alison Hills, in 'Understanding Why' (2016), holds that understanding why $p$ because $q$ is a cognitive achievement constituted by a cluster of abilities she calls cognitive control: being able to follow an explanation given by another, to give it in one's own words, to draw the conclusion $p$ from the information that $q$, and — the crucial one for the present argument — to do the corresponding thing for cases relevantly similar to the original, moving from $q^*$ to $p^*$. That last ability is deformation under another name. Hills argues that these a