Population Genetics Under Dynamic Demography — Epoche C1
A population of 100 breeding adults, of which 10 are male and 90 female, loses genetic variation at the rate a population of 36 would. That number is its effective population size , written $N_e$: not a head count but the size of an idealised population that would drift as fast as the real one does. The gap between 100 and 36 is the whole subject of this essay, because $N_e$ enters almost every quantitative prediction in population genetics, and because the classical treatment of it as a constant is the assumption that dynamic demography breaks. The compressed version of this essay called for $N_e$ to be modelled as a stochastic rather than a fixed quantity. That call is right in spirit and needs three corrections in detail: one of its claims about the harmonic mean is the wrong way round; the proposal that $N_e$ be treated as a random variable is not quite the right formalisation of the problem; and the research programme it proposes has in substantial part already been carried out. Working through the derivations shows where each correction bites. Where $N_e$ comes from, and why it is not a head count The reference model is Wright–Fisher: a diploid population of constant size $N$, so $2N$ gene copies, in which each generation is formed by drawing $2N$ copies at random with replacement from the previous generation's $2N$. Reproduction is otherwise unstructured — no selection, no mating system, no age classes. Drift falls out immediately. Take two gene copies in the offspring generation; the chance that they were drawn from the same parental copy is $1/(2N)$, since there are $2N$ copies to choose from. If they came from the same copy they are necessarily identical; if not, they are identical only if the two parental copies already were. So the heterozygosity $H$ — the probability that two randomly drawn copies differ — is multiplied by $\bigl(1 - 1/(2N)\bigr)$ each generation, giving $$H_t = H_0\left(1 - \frac{1}{2N_e}\right)^{t}.$$ The formula defines $N_e$ rather than merely using it: $N_e$ is the value that makes the ideal model reproduce the real population's rate of loss. Two standard corrections show how far that value can sit from the census size $N$. Unequal sex ratio is the first. Every offspring takes half its genes from a male and half from a female, so the male pool and the female pool each contribute half the genes regardless of how many individuals are in them. Wright's 1931 result is $$N_e = \frac{4 N_m N_f}{N_m + N_f},$$ which for $N_m = 10$ and $N_f = 90$ gives $4 \times 10 \times 90 / 100 = 36$: the opening number. The scarce sex is a bottleneck through which half the genome must pass every generation. Variance in reproductive success is the second, and it is the one the original essay invoked without a formula. Let $V_k$ be the variance in the number of surviving offspring per individual, with the mean fixed at $\bar k = 2$ so that the population holds steady. Crow and Kimura (1970) give $$N_e = \frac{4N - 2}{V_k + 2}.$$ Under Wright–Fisher reproduction the offspring count is Poisson, so $V_k = \bar k = 2$ and $N_e = (4N-2)/4 \approx N$, as it must be. Raising $V_k$ to 6 — a few individuals monopolising reproduction — gives $N_e \approx N/2$. Lowering it towards 0, which is what strong competition for a limited number of breeding slots does, drives $N_e$ towards $2N - 1$. That is the ceiling: no reproductive regime can make a population drift more slowly than twice its census size, because with $V_k=0$ each individual contributes exactly two offspring and the only remaining randomness is Mendelian segregation. This substantiates the original essay's claim that competition at high density can push $N_e$ above census size, and bounds it. Why averaging over time is harmonic, not arithmetic Now let size vary across generations, $N_1, N_2, \ldots, N_t$. The retention of heterozygosity multiplies across generations: $$H_t = H_0 \prod_{i=1}^{t}\left(1 - \frac{1}{2N_i}\right).$$ Taking logarithms and using $\ln(1-x) \approx -x$ for the small quantities $1/(2N_i)$ turns the product into a sum, $\ln(H_t/H_0) \approx -\sum_i 1/(2N_i)$. Setting this equal to the constant-size expression $-t/(2N_e)$ and cancelling gives the formula the original essay quoted: $$\frac{1}{N_e} = \frac{1}{t}\sum_{i=1}^{t}\frac{1}{N_i}.$$ The average is harmonic because what accumulates additively is the rate of drift, and the rate is $1/(2N)$ per generation, not $N$. Reciprocals add; sizes do not. The consequence is severe. Take ten generations, nine at $N = 1000$ and one at $N = 10$. The arithmetic mean is 901. The harmonic calculation gives $\frac{1}{N_e} = \frac{1}{10}\left(\frac{9}{1000} + \frac{1}{10}\right) = 0.0109$, so $N_e = 92$. One generation in ten spent at ten individuals reduces the effective size by a factor of about ten, because that single generation contributes $0.1$ to a sum whose other nine terms total $0.009$. Here a claim in the original text must be corrected. It said that the harmonic mean, being dominated by bottlenecks, risks "underestimating the overall long-term genetic diversity" when bottlenecks are rare but severe. The harmonic mean does not underestimate anything of the sort: the derivation above shows it is the exact leading-order answer for the cumulative decay of expected heterozygosity, bottlenecks included. Its real limitation is different and more interesting. Expected heterozygosity is a single summary statistic, and a great many demographic histories share the same one. A population that was small once and large ever since, and a population that was always of size 92, agree on $H_t$ and disagree completely on the shape of the genealogy and hence on the site frequency spectrum — the distribution of how many sampled sequences carry each variant. Collapsing a history to one number discards precisely the information that distinguishes histories. The correction that matters most: there is no single $N_e$ The original essay treats $N_e$ as one quantity that happens to