Linkage Disequilibrium and the Binomial Variance of Allele Frequencies — Epoche C1
The claim, and the binomial it is about Alleles at nearby positions on a chromosome are inherited together, so their frequencies in a sample are correlated — and a widely repeated inference from this fact is that allele frequencies must therefore fluctuate more than the Wright–Fisher model predicts, blurring the signals that population geneticists use to detect natural selection. This essay argues that the inference is misdirected. The correlation is real and consequential, but it does not touch the quantity it is usually accused of inflating, and locating its effect correctly changes what an observed excess of variance should be taken to mean. The model at issue is the Wright–Fisher model, the standard null hypothesis of the field: a population of constant size $N$ diploid individuals, non-overlapping generations, random mating, and no selection, mutation or migration. Each of the $2N$ gene copies of the next generation is drawn independently, with replacement, from the gene pool of the present one. If an allele has frequency $p$ now, its count $k$ next generation is therefore binomially distributed, $k \sim B(2N, p)$, and since a binomial count has variance $2Np(1-p)$, the frequency $k/(2N)$ has variance $$\operatorname{Var}(\Delta p) = \frac{p(1-p)}{2N}.$$ The factor $p(1-p)$ is there because sampling error vanishes when the allele is absent or universal; the division by $2N$ is there because each individual contributes two gene copies, so a population of $N$ individuals supplies $2N$ independent draws. This is what "genetic drift" means quantitatively, and it is the yardstick against which any claim of overdispersion — variance in excess of a model's prediction — must be measured. Linkage disequilibrium , abbreviated LD, is the non-random association of alleles at different loci: if allele $A$ at one locus and allele $B$ at another occur together on the same chromosome more often than their separate frequencies would imply, the two loci are in disequilibrium. The standard measure begins with the coefficient $D = P_{AB} - P_A P_B$, the excess of the joint haplotype frequency over the product of the marginal allele frequencies, and normalises it: $$r^2 = \frac{(P_{AB} - P_A P_B)^2}{P_A(1-P_A)\,P_B(1-P_B)}.$$ The reason this particular normalisation is used, and the fact that makes the rest of the argument work, is that $r^2$ is exactly the squared Pearson correlation coefficient between two indicator variables — one recording whether a randomly chosen chromosome carries $A$, the other whether it carries $B$. The denominator is the product of their variances, since an indicator with mean $P_A$ has variance $P_A(1-P_A)$. So LD is a correlation, and the question of what it does to variance is a question about what correlations do to variances. Why the binomial survives at a single locus That question has a short answer at one locus, and it is the core of the essay's thesis. Consider the sampling step of the Wright–Fisher model with many loci at once. The gametes of the next generation are drawn independently from the current pool of haplotypes — whole chromosomes, with all their loci — so the joint distribution of the next generation's haplotype counts is multinomial, and LD is precisely the statement that the haplotype frequencies are not products of allele frequencies. Now ask about one locus. A drawn gamete carries allele $A$ with probability $P_A$, which is the total frequency of all haplotypes containing $A$, and the draws are independent across gametes. The count of $A$ is therefore exactly $B(2N, P_A)$, whatever the haplotype frequencies were. Marginalising a multinomial gives a binomial, and it does so regardless of the correlations among the categories. So the per-locus null distribution is untouched by LD, exactly rather than approximately. This is stronger than the version of the claim the compressed argument makes — that the frequency at a locus "still tends towards a binomial distribution if the sample size is sufficiently large". No largeness is needed; the binomial is exact under the model's assumptions. What can break it is any departure from independent sampling of gene copies: a few families contributing disproportionately many offspring, selfing, or a recent bottleneck. Those reduce the effective population size $N_e$, the size of an idealised population that would drift at the observed rate, and thereby inflate $p(1-p)/(2N_e)$ relative to $p(1-p)/(2N)$. They are, however, facts about who reproduces, not about which alleles travel together. What linkage disequilibrium does inflate The previous section vindicates the essay's thesis at one locus, and it is here that the thesis needs a correction rather than a defence, because LD is not a small effect misplaced — it is a large effect misplaced. Correlations do not change marginal variances, but they change the variance of averages, which is where nearly every genome-wide analysis lives. Let $Z_1, \dots, Z_L$ be per-locus statistics, each with variance $v$ and average pairwise correlation $\bar{\rho}$. Their mean has variance $$\operatorname{Var}\!\left(\frac{1}{L}\sum_{i=1}^{L} Z_i\right) = \frac{v}{L}\Big(1 + (L-1)\bar{\rho}\Big) = \frac{v}{L_{\text{eff}}}, \qquad L_{\text{eff}} = \frac{L}{1 + (L-1)\bar{\rho}},$$ the first equality following by expanding the square and counting the $L$ variance terms and the $L(L-1)$ covariance terms. The quantity $L_{\text{eff}}$ is the number of independent loci that would give the same precision. The dangerous feature of the formula is the product $L\bar{\rho}$: with $L = 10^5$ loci and an average correlation of only $0.01$, the bracket is about $1000$ and $L_{\text{eff}}$ is about $100$. A correlation far too small to notice between any particular pair of loci reduces the information in a hundred thousand of them to that of a hundred. How large is $\bar{\rho}$ in practice? Under drift and recombination alone, the expected squared correlation between two loci separated by recombination