The Underappreciated Role of Fluctuating Selection — Epoche C1
The quantity in question, and what it is measured against Effective population size, written $N_e$, is the number that answers the question "how fast do allele frequencies wander by chance in this population?" — and it is almost never equal to the number of individuals one could count. It is defined by comparison: $N_e$ is the size of an idealised population, one with random mating, non-overlapping generations, equal sex ratio and Poisson-distributed offspring numbers, that would lose genetic variation at the same rate as the real population under study. An allele is a variant form of a gene at a locus, and its frequency $p$ in the population drifts up and down from one generation to the next simply because only a finite sample of gametes founds the next generation; $N_e$ calibrates the size of that wandering. Two of the classical formulas are worth deriving rather than quoting, because the derivations show what kind of thing $N_e$ is. Consider first a population with $N_m$ breeding males and $N_f$ breeding females. Take two gene copies at random from the offspring generation and ask for the probability that they descend from the same copy in the parental generation — the coalescence probability, which for a diploid population of effective size $N_e$ is by definition $1/(2N_e)$. Each of the two copies is paternal or maternal with probability $\tfrac{1}{2}$. Both are paternal with probability $\tfrac{1}{4}$; they then came from the same father with probability $1/N_m$, and given that, from the same one of his two gene copies with probability $\tfrac{1}{2}$. The maternal case is parallel, and a paternal copy cannot coalesce with a maternal one in a single generation. Hence $$\frac{1}{2N_e} = \frac{1}{8N_m} + \frac{1}{8N_f}, \qquad \text{so} \qquad N_e = \frac{4 N_m N_f}{N_m + N_f}.$$ The factor $4$ is the product of the two halves that were spent on choosing a sex and then a gene copy within a parent. With one breeding male and ninety-nine breeding females, a census of $100$ yields $N_e = 4 \times 99 / 100 = 3.96$: the population drifts as though it contained four individuals. The second formula, for variation in family size, was misstated in the original version of this note and should be corrected. It was given as $N_e = N_c/(1 + V_k/\bar{k})$, where $\bar{k}$ is the mean number of offspring per individual and $V_k$ the variance in that number. Every formula for $N_e$ must pass one test: applied to the idealised population itself, it must return $N_e = N$. In the idealised population offspring numbers are Poisson, so $V_k = \bar{k}$, and for constant census size $\bar{k} = 2$; the formula as printed then gives $N_e = N/2$, which is wrong by a factor of two. Wright's result for a diploid population of constant size is $$N_e = \frac{4N - 2}{V_k + 2},$$ which returns $N_e = N - \tfrac{1}{2}$ under the Poisson check, as it must. The formula makes the biology visible: a species in which a few individuals monopolise reproduction, say $V_k = 10$, has $N_e \approx 4N/12 = N/3$. Now the point that bears directly on this essay's thesis. Suppose the census size itself varies across generations, taking values $N_1, \ldots, N_t$. Two lineages fail to coalesce over $t$ generations with probability $\prod_i (1 - 1/(2N_i)) \approx \exp\!\left(-\sum_i 1/(2N_i)\right)$, and equating this to $\exp(-t/(2N_e))$ gives $$\frac{1}{N_e} = \frac{1}{t}\sum_{i=1}^{t} \frac{1}{N_i},$$ the harmonic mean, which is dominated by the smallest terms. Sizes of $1000$, $1000$ and $10$ in three successive generations give $N_e = 3/(0.001 + 0.001 + 0.1) \approx 29$, not $670$. Environmental variability, in other words, is already inside the classical framework — provided it acts on numbers. The gap the present essay identifies is elsewhere. What $N_e$ is silent about Every quantity above was computed by tracing ancestry with no reference to fitness. That is not an oversight but the definition: $N_e$ is a neutral-locus construct, describing the rate of sampling noise at a site whose variants have no effect on survival or reproduction. It does not assume that selection is constant, as the original text put it; it makes no statement about selection at all. What it supplies is a yardstick against which selection is measured. The yardstick is the product $N_e s$, where $s$ is the selection coefficient — the proportional fitness advantage of one allele over another, so that carriers of the favoured type leave $1+s$ times as many descendants. When $|N_e s| \ll 1$ the deterministic push is smaller than the generation-to-generation noise and the allele behaves as though neutral; when $N_e s \gg 1$ selection dominates. Haldane's branching-process argument gives the sharpest version for a single new beneficial mutation: treating its early spread as a family tree in which each copy leaves a Poisson number of copies with mean $1 + s$, the probability that the lineage escapes extinction rather than being lost by chance in its first few generations is approximately $2s$. A mutation with a one per cent advantage is lost, despite being advantageous, ninety-eight times out of a hundred. All of this presumes a fixed $s$. The question this essay raises is what happens when $s$ is not fixed but is itself drawn afresh each generation from the environment. Using the essay's own equation properly The change in the frequency of an allele under genic selection, with the favoured type having relative fitness $1+s$ and the alternative $1$, is $$\Delta p = p(1-p)\,\frac{s}{1 + sp},$$ the denominator being the mean fitness of the population, which normalises the frequencies. The original note wrote this equation and then left it. Expanding it is what makes the role of variance appear. For small $s$, $\frac{s}{1+sp} = s - s^2 p + O(s^3)$, so if $s$ is a random variable with mean $\bar{s}$ and variance $\sigma_s^2$, the expected change per generation is $$\mathbb{E}[\Delta p] = p(1-p)\left(\bar{s} - \overline{s^2}\,p\right) \approx p(1-p)\left(\bar{s} - \si