Unravelling Gut Microbiome Complexity through Network and Causal Inference Models — Epoche C2
What a diversity index can and cannot see The claim that a gut community with a higher Shannon index is a healthier one is the working assumption behind a large fraction of published 16S rRNA gene surveys, and it fails for a reason that can be stated as a theorem rather than a suspicion. The two indices in question are standard. The Shannon index, $H = -\sum_{i=1}^{S} p_i \ln(p_i)$, is the entropy of the abundance distribution, where $p_i$ is the proportional abundance of taxon $i$ and $S$ the number of taxa observed; it quantifies the uncertainty in predicting the identity of one individual drawn at random. The Simpson index, written here as $D = 1 - \sum_{i=1}^{S} p_i^2$, is the probability that two individuals drawn at random, with replacement, belong to different taxa. Both are correct as far as they go, and both are members of one family. That family is the Hill numbers, which are worth putting on the page because they make the defect visible. For an order parameter $q \geq 0$, $$^{q}\!D = \left(\sum_{i=1}^{S} p_i^{q}\right)^{1/(1-q)},$$ with $q = 0$ giving the richness $S$, the limit $q \to 1$ giving $\exp(H)$, and $q = 2$ giving $1/\sum_i p_i^2$, the inverse Simpson concentration. Each is an effective number of taxa: the number of equally abundant taxa a community would need in order to score as it does. The order $q$ is a weighting knob, and raising it shifts weight from rare taxa towards common ones. One consequence, which Hill set out in 1973, is that the Shannon entropy is not itself a diversity but the logarithm of one, so that differences in $H$ between samples are not proportional to differences in the quantity anyone means. The deeper consequence is structural. Every Hill number is a symmetric function of the vector $(p_1, \dots, p_S)$: it is invariant under any permutation of the taxon labels. Two communities with identical abundance profiles but entirely disjoint membership therefore receive identical scores at every order $q$, and no amount of care in estimating $H$ changes this. It follows immediately that no diversity index can encode which taxa are present, and a fortiori that none can encode how they interact. This is worth stating plainly because the earlier version of this essay located the flaw elsewhere, in "an underlying assumption of uniform stochasticity in species interactions". The indices make no assumption about interactions at all. They are functions of the abundance vector and nothing else, and that — not a hidden and false ecological premise — is why they are silent about structure. A second, less obvious point: a single index does not even fix the ordering between two communities. Consider community C with five taxa at proportions $(0.6, 0.1, 0.1, 0.1, 0.1)$ and community D with three taxa at $(0.4, 0.4, 0.2)$. For C, $H = 1.228$ nats, so $\exp(H) = 3.41$ effective taxa, while $\sum_i p_i^2 = 0.40$ and $^{2}\!D = 2.50$. For D, $H = 1.055$ nats, so $\exp(H) = 2.87$, while $\sum_i p_i^2 = 0.36$ and $^{2}\!D = 2.78$. At $q = 1$ community C is the more diverse; at $q = 2$ community D is. The ranking reverses because C's diversity is carried by four rare taxa that the $q = 1$ weighting rewards and the $q = 2$ weighting nearly ignores. "Higher diversity" is therefore not a well-formed comparison until the order is specified. The same arithmetic explains why an index cannot register a keystone taxon — a taxon whose influence on community structure and function is disproportionate to its abundance, in the sense used by Banerjee, Schlaeppi and van der Heijden in their 2018 review, where such taxa are identified by their position in an inferred network rather than by how much of the community they constitute. A taxon at $0.1$ per cent relative abundance contributes $-0.001 \ln(0.001) = 0.0069$ nats to a Shannon index whose typical value in a gut survey is of order $3.5$ nats: about $0.2$ per cent of the total. Its contribution to $\sum_i p_i^2$ is $10^{-6}$, which is beyond the resolution of any realistic sampling. Whatever it does to the community, the index cannot see it happen. This is what gives the earlier version's thought experiment its force. Two individuals with the same Shannon score may share not a single taxon; one may carry a well-connected guild of short-chain fatty acid producers and mucin degraders and the other a fragmented assembly of opportunists. The indices are identical by construction. Everything that distinguishes the two cases lives in information the indices discard. Diversity and stability: what the only general theorem says Before turning to what should replace the indices, the assumption they are used to support deserves direct examination, because it is not merely unsupported but contradicted by the one general result available. The assumption is that greater richness confers functional robustness. Robert May's 1972 argument concerns a community of $S$ species near an equilibrium, described by a community matrix $A$ whose entry $A_{ij}$ is the effect of species $j$ on the growth rate of species $i$ at that equilibrium. Set the diagonal to $-1$, so each species is self-regulating on a unit timescale. Let each off-diagonal entry be non-zero with probability $C$, the connectance, and when non-zero drawn independently with mean zero and variance $\alpha^2$. As $S$ grows, the eigenvalues of the random off-diagonal part fill a disc in the complex plane of radius $\alpha\sqrt{SC}$; adding the diagonal shifts the whole spectrum left by one, so the rightmost eigenvalue sits near $-1 + \alpha\sqrt{SC}$. Local stability requires that this be negative, giving May's criterion $$\alpha\sqrt{SC} The arithmetic runs against the received wisdom. At $S = 100$ taxa and $C = 0.1$, stability requires $\alpha The hypotheses are where the argument for the microbiome must be made, and it is a real dispute rather than a settled one. May's entries are independent and structureless; real microbial interactions are not, since cross-feeding pairs have correlat