The Wright-Fisher Model's Distortions in Structured Populations — Epoche C1
The shortcut this essay examines A population geneticist modelling a species that breeds continuously — a bird that lives eight years and lays every spring, an oak that sheds acorns for two centuries — is usually advised to use the Wright–Fisher model anyway and to compensate by rescaling: substitute an effective population size for the census number, measure time in generations rather than years, and proceed. That advice is not folklore. It is the informal statement of a limit theorem, and like any limit theorem it holds only under hypotheses. This essay states the theorem, shows that overlapping generations and age structure satisfy its hypotheses closely enough for the rescaling to be legitimate rather than merely convenient, and locates the places where it genuinely fails. Two claims in the compressed version are corrected along the way: it overstated the damage overlapping generations do to a genealogy's shape, and it missed a theorem of Maruyama's limiting the damage spatial structure does to a beneficial mutation's fixation probability. The model stated exactly, and a convention that needs fixing Take a haploid population of $N$ gene copies at one locus with two alleles, $A_1$ at frequency $p$ and $A_2$ at frequency $1-p$. Each copy in the next generation independently picks a parent uniformly at random from the current $N$ and inherits its allele, so the number $K$ of $A_1$ copies next generation is binomial, $$P(K = k \mid p) = \binom{N}{k}\, p^{k}\,(1-p)^{N-k},$$ so that $E[K/N] = p$ — no systematic change — and $\operatorname{Var}(K/N) = p(1-p)/N$. Drift, in this model, is nothing but the sampling variance of a binomial draw. Two things in the compressed version need correcting here. That formula is the neutral case, yet it was introduced with the remark that parental contributions are determined by fitness, which it does not represent. Selection enters by replacing the sampling probability with the post-selection frequency: if $A_1$ carries relative fitness $1+s$, a parent is drawn from a pool in which $A_1$ has frequency $$p^{*} = \frac{p(1+s)}{p(1+s) + (1-p)} = \frac{p(1+s)}{1+sp},$$ and $K \sim \text{Binomial}(N, p^{*})$. Second, the compressed version wrote $\binom{N}{k}$ for a population described as having $N$ individuals, then quoted a pairwise coalescence rate of $1/(2N_e)$. Those are two different bookkeeping conventions. In a diploid population of $N$ individuals there are $2N$ gene copies at an autosomal locus, so the coefficient is $\binom{2N}{k}$ and the exponent $2N-k$. This is not pedantry: every timescale below inherits the factor of two, and mixing the conventions is the commonest way of getting an effective size wrong by exactly that factor. What the rescaling shortcut actually is The justification for replacing a real population by a Wright–Fisher one is a theorem about genealogies, so the genealogical picture comes first. Trace two distinct gene copies backwards. Each chose its parent uniformly and independently, so in any one generation they have the same parent — they coalesce , in the standard term — with probability $1/N$ in the haploid model, or $1/(2N)$ for a diploid population of $N$ individuals. The number of generations back to their common ancestor is geometric with that success probability, mean $2N$ in the diploid case; measure time in units of $2N$ generations and it converges to an exponential of rate one. For a sample of $n$ copies the same argument gives each of the $\binom{n}{2}$ pairs an independent rate of one, while the probability that three or more copies share a parent in the same generation is of order $N^{-2}$ per generation and so vanishes once time is stretched by a factor of $N$. The limiting object — a random binary tree in which $k$ surviving lineages wait an exponential time of rate $\binom{k}{2}$ and then two, chosen uniformly, merge — is Kingman's $n$-coalescent (Kingman 1982). The compressed version called it the $N$-coalescent; the index is the sample size, and the whole point of the construction is that the population size has been scaled away. Now the theorem. Drop Wright–Fisher sampling and require only that the offspring numbers $\nu_1,\dots,\nu_N$ contributed by the $N$ parents are exchangeable — their joint distribution unchanged by relabelling the parents — and sum to $N$, so population size is constant. Exchangeability forces $E[\nu_1]=1$; let $\sigma_N^2 = \operatorname{Var}(\nu_1)$. Pick two distinct offspring at random; they share a parent with probability $$\frac{\sum_i E[\nu_i(\nu_i-1)]}{N(N-1)} = \frac{E[\nu_1(\nu_1-1)]}{N-1} = \frac{E[\nu_1^2] - E[\nu_1]}{N-1} = \frac{\sigma_N^2}{N-1},$$ using $E[\nu_1^2] = \sigma_N^2 + 1$ at the last step. Kingman proved that if $\sigma_N^2$ converges to a finite positive $\sigma^2$, and no single parent can claim a non-vanishing fraction of the offspring (a third-moment condition), then such a model's sample genealogy, with time in units of $N/\sigma^2$ generations, converges to the very same $n$-coalescent. Under Wright–Fisher sampling $\nu_1$ is $\text{Binomial}(N,1/N)$, so $\sigma_N^2 \to 1$ and the time unit is $N$ generations, as computed above. This is the content of the shortcut, and it is worth stating baldly. The entire difference between a Wright–Fisher population and any other exchangeable one is a single number, the variance in offspring number, absorbed into $N_e = N/\sigma^2$; nothing else about the reproductive biology survives into the limit. Correspondingly, the ways a real population can distort the model are exactly the ways it can violate the three hypotheses. Overlapping generations: the shortcut survives intact The cleanest test case is the Moran model (Moran 1958), which removes discrete generations altogether. At each step one individual chosen uniformly at random reproduces, its single offspring enters the population, and one individual chosen uniformly at random dies. Generations overlap completely: a newborn competes with individuals of every ag