The Enduring Role of Frequency-Dependent Selection in Genetic Diversity — Epoche C1
A fish with a crooked mouth, and a distinction that has to be made first The scale-eating cichlid Perissodus microlepis of Lake Tanganyika comes in two forms: one whose mouth is twisted to the left, which attacks its prey's right flank, and one twisted to the right, which attacks the left. Michio Hori reported in 1993 that the proportion of the two morphs in the lake oscillated about equality over a run of years, and gave the mechanism: prey fish learn to guard the side they are most often attacked from, so whichever morph is currently in the minority feeds more successfully and increases, until it becomes the majority and the advantage passes to the other. This is negative frequency-dependent selection — selection in which a type's fitness declines as it becomes more common — and it is the clearest kind of evidence for the argument this essay makes: that mechanisms of this sort maintain genetic variation that neither genetic drift nor mutation would maintain, and that they do so through a dynamic which purely neutral models cannot produce. Before the argument can be developed, however, an error in the compressed version of this essay has to be corrected, because it concerns the classification of its own leading example. That version introduced heterozygote advantage — the case in which the heterozygote $A_1A_2$ is fitter than either homozygote, also called overdominance — as 'a classic example of frequency-dependent selection'. It is not one. Under overdominance the fitnesses of the three genotypes are fixed numbers that do not depend on how common anything is; what varies with frequency is the average fitness of an allele , and it varies because the mix of genotypes an allele finds itself in depends on frequency. The two mechanisms produce superficially similar outcomes and are structurally quite different, and the difference is the subject of the next two sections. The selection recursion, derived Everything below rests on one recursion, which is worth deriving rather than quoting. Consider a single locus with two alleles, $A_1$ at frequency $p$ and $A_2$ at frequency $q = 1-p$, in a large randomly mating population, so that the genotypes occur in the proportions $p^2$, $2pq$ and $q^2$. Let the three genotypes have fitnesses $w_{11}$, $w_{12}$ and $w_{22}$, where fitness means expected relative contribution of offspring to the next generation. An $A_1$ allele sits in an $A_1A_1$ homozygote with probability $p$ and in a heterozygote with probability $q$, since its partner allele is drawn at random from the population. Its expected fitness — its marginal fitness — is therefore $$w_1 = p\,w_{11} + q\,w_{12}, \qquad w_2 = p\,w_{12} + q\,w_{22},$$ and the population mean fitness is $\bar{w} = p\,w_1 + q\,w_2$. Selection reweights the alleles in proportion to these marginal fitnesses, so $p' = p\,w_1/\bar{w}$, and $$\Delta p = p\,\frac{w_1 - \bar{w}}{\bar{w}} = \frac{p\,q\,(w_1 - w_2)}{\bar{w}},$$ the last step following because $w_1 - \bar{w} = w_1 - p w_1 - q w_2 = q(w_1 - w_2)$. This is the equation the compressed version gave, and it is correct. Its structure explains the whole subject: the direction of change is set by the sign of $w_1 - w_2$, and the magnitude is throttled by the factor $pq$, which vanishes at both ends, so change is slowest when one allele is rare. Overdominance: the equilibrium, and why it is not frequency-dependent Now put in the overdominant fitnesses. Take $w_{11} = 1-s_1$, $w_{12} = 1$ and $w_{22} = 1-s_2$, with $s_1, s_2 > 0$, so the heterozygote is fittest and $s_1$ and $s_2$ measure how much each homozygote falls short of it. Then $$w_1 - w_2 = \big[p(1-s_1) + q\big] - \big[p + q(1-s_2)\big] = q\,s_2 - p\,s_1 .$$ Setting this to zero and substituting $q = 1-p$ gives $s_2 - p s_2 = p s_1$, hence $$\hat{p} = \frac{s_2}{s_1 + s_2}.$$ The equilibrium favours whichever allele has the fitter homozygote, as it should: the numerator is the deficit of the other homozygote. It is stable, and the reason is legible in the expression above rather than assumed: $q s_2 - p s_1$ is positive when $p$ is below $\hat p$ and negative when $p$ is above it, so $\Delta p$ always points back. The mechanism is exactly the one the compressed version described, and it can be quantified, which is what shows it to be something other than frequency-dependent selection. Of all the $A_1$ copies in the population, the fraction carried by heterozygotes is $$\frac{2pq}{2p^2 + 2pq} = q .$$ As $A_1$ becomes rare, $q$ approaches one and essentially every copy of $A_1$ sits in a heterozygote, so its marginal fitness approaches $w_{12} = 1$ while the common allele's marginal fitness approaches $w_{22} = 1-s_2$. The rare allele enjoys an advantage of $s_2$ purely because rarity and heterozygosity coincide under random mating. No genotype's fitness changed; the census of genotypes did. The canonical instance is the sickle-cell haemoglobin variant, whose heterozygote carriers Allison showed in 1954 to be protected against falciparum malaria, while the homozygote suffers sickle-cell disease: three fixed fitnesses, one environment, and a polymorphism maintained by their ordering. What genuine frequency dependence adds The turn in the argument is here. Under frequency-dependent selection the genotypic fitnesses are themselves functions of the frequencies, so the fitness parameters that were constants above become variables. The simplest case that shows the difference is a haploid one, where the heterozygote route to polymorphism is unavailable by construction. Let two types have fitnesses $$w_1 = 1 - c\,p, \qquad w_2 = 1 - c\,q,$$ with $c > 0$ measuring how much a type suffers from its own abundance — through competition for a shared resource, or through a predator that has learned its appearance. Then $w_1 - w_2 = c(q - p) = c(1-2p)$, and the recursion becomes $\Delta p \propto p\,q\,c\,(1-2p)$, giving a stable equilibrium at $\hat p = 1/2$. Polymorphism is maintained with no dominance, no heterozygotes