The Stochastic Nature of Allele Frequencies in Finite Populations — Epoche B2
Beyond Determinism: The Stochastic Nature of Allele Frequencies in Finite Populations The Hardy-Weinberg principle stands as a cornerstone in population genetics, providing a fundamental null model against which real-world evolutionary changes can be measured. It describes a theoretical population where allele and genotype frequencies remain constant across generations, undisturbed by evolutionary forces. This principle, however, rests on a set of idealised assumptions, one of the most critical being an infinitely large population size. Selection, mutation and gene flow are the well-recognised reasons why allele frequencies fail to stay constant, and non-random mating the well-recognised reason why genotype proportions depart from $p^2 : 2pq : q^2$; the stochasticity that comes from a population merely being finite—genetic drift—is a fourth kind of departure [1] , and it is regularly confused with the other three. This essay will explore how genetic drift fundamentally violates the assumptions of the Hardy-Weinberg model, demonstrating its predictable role in the decay of genetic diversity over time and challenging the view that selection is always the primary driver of allele frequency change. The Hardy-Weinberg Principle: A Deterministic Null Model The Hardy-Weinberg principle describes the relationship between allele frequencies and genotype frequencies in an idealised, non-evolving population. Consider a single genetic locus with two alleles, $A$ and $a$. Let $p$ represent the frequency of allele $A$ and $q$ represent the frequency of allele $a$ in the population's gene pool. Since these are the only two alleles at this locus, their frequencies must sum to one: $$ p + q = 1 $$ Under the principle's assumptions, if individuals mate randomly, the probability of forming a diploid genotype is simply the product of the probabilities of drawing the constituent alleles from the gene pool. For instance, the probability of an individual inheriting two $A$ alleles (genotype $AA$) is $p \times p = p^2$. Similarly, the probability of inheriting two $a$ alleles (genotype $aa$) is $q \times q = q^2$. For the heterozygous genotype $Aa$, there are two ways to form it: inheriting $A$ from one parent and $a$ from the other, or vice versa. Thus, the frequency of $Aa$ is $p \times q + q \times p = 2pq$. The sum of these genotype frequencies must also equal one: $$ p^2 + 2pq + q^2 = 1 $$ This equation represents the Hardy-Weinberg equilibrium. For these frequencies to remain constant across generations, the model makes five crucial assumptions, in the form given by Griffiths and colleagues: No mutation: No new alleles are introduced, and existing alleles do not change. No gene flow: There is no migration of individuals into or out of the population. Random mating: Individuals choose mates without regard to their genotype. No natural selection: All genotypes have equal survival and reproductive rates. Infinitely large population size: This assumption ensures that allele frequencies are not subject to random fluctuations due to chance events. These four are not four versions of the same thing, and separating them matters for everything that follows. Mutation, gene flow and selection change the allele frequencies $p$ and $q$. Non-random mating does not: it changes only the genotype proportions into which a fixed pair of allele frequencies is packaged. The cleanest case is complete selfing, where $AA$ and $aa$ breed true and $Aa$ gives $\tfrac{1}{4}AA : \tfrac{1}{2}Aa : \tfrac{1}{4}aa$, so that heterozygotes are halved every generation while not a single allele copy is gained or lost: $$ f_{Aa}^{(t+1)} = \tfrac{1}{2}\,f_{Aa}^{(t)}, \qquad p^{(t+1)} = f_{AA}^{(t+1)} + \tfrac{1}{2}f_{Aa}^{(t+1)} = p^{(t)} $$ Keep that distinction in view: a departure from the proportions $p^2 : 2pq : q^2$ within one generation is a statement about mating and packaging, whereas a change in $p$ between generations is a statement about evolutionary force. The fifth assumption is of a different kind again. It is not a force at all but a limit, and it is what makes the model deterministic; in reality no population is infinitely large, and the violation of this assumption gives rise to genetic drift. Genetic Drift: Random Sampling in Finite Populations In any real population, the number of individuals is finite. This means that the alleles passed from one generation to the next are not a perfect, infinite sample of the parental gene pool, but rather a finite, random sample. This process of random sampling introduces an element of chance into allele frequency changes, a phenomenon termed genetic drift . Unlike selection, which causes directional changes in allele frequencies based on fitness differences, genetic drift causes non-directional, random fluctuations. To understand genetic drift quantitatively, we often employ the Wright-Fisher model [2] , an idealised mathematical framework for population genetics. This model simplifies the complexities of real populations to focus on the effect of random sampling. Its key assumptions are: Discrete generations: Reproduction occurs in distinct, non-overlapping generations. Fixed population size ($N$): The number of diploid individuals in the population remains constant. Random mating: Any individual can mate with any other individual. No mutation, selection, or gene flow: These deterministic forces are explicitly excluded to isolate the effect of drift. All individuals contribute equally to a gamete pool: In each generation, all $2N$ alleles (from $N$ diploid individuals) contribute to a single, large gamete pool. The next generation's $2N$ alleles are then formed by randomly drawing from this pool. Under the Wright-Fisher model, if allele $A$ has a frequency $p$ in the current generation, the number of $A$ alleles in the next generation, let's call it $K_A$, is determined by drawing $2N$ alleles from the gamete pool. This process precisely follows a binomial distribution. The probability