Information Flow in Eukaryotic Gene Expression — Epoche B2
Beyond the Central Dogma: Information Flow in Eukaryotic Gene Expression The Central Dogma is usually recited as an arrow chain [1] , $\text{DNA} \to \text{RNA} \to \text{protein}$, and read as the claim that this pathway runs in one direction only. That is not what Francis Crick stated. His dogma, put forward in 1958 and restated in 1970, concerns the residue-by-residue transfer of sequence information, and its content is a prohibition: once information has passed into protein it cannot get out again, either to another protein or to a nucleic acid. Transfers among nucleic acids are expressly permitted, which is why the discovery of reverse transcription in 1970 did not overturn the dogma — a point Crick made himself. Nothing in this essay disturbs that statement. What the arrow chain silently invites, and what this essay goes beyond, is a different claim: that a gene's sequence fixes how much of its protein a cell will contain. It does not. Genetically identical cells, in the same medium, at the same point in the cell cycle, contain measurably different numbers of copies of the same protein, and the spread is not small. Eukaryotic expression makes this worse rather than better, because transcription and translation are separated by a nuclear membrane and an export step, and because promoters spend long periods in chromatin states that polymerase cannot read. This essay dissects that cell-to-cell variability into the part arising from the stochastic mechanics of a single gene (intrinsic [2] ) and the part arising from fluctuations in the shared cellular environment (extrinsic), shows how a single experiment separates them, and closes by asking what the separation means in a eukaryote, where the experiment was not originally done. Quantifying Variability: The Coefficient of Variation Gene expression is a series of biochemical reactions involving a finite and often small number of molecules: messenger RNA (mRNA) transcripts, polymerases, ribosomes, proteins. Because each reaction occurs probabilistically, the output is a random variable rather than a number, and identical cells in identical conditions receive different draws from it. Write $P$ for the protein copy number in a cell. Across a population of $N$ cells with measured levels $P_i$, the mean is $$ \langle P \rangle = \frac{1}{N} \sum_{i=1}^N P_i $$ and the spread about it is the variance, $$ \text{Var}(P) = \frac{1}{N-1} \sum_{i=1}^N (P_i - \langle P \rangle)^2 $$ Variance alone cannot be compared between genes, because it carries the units of $P$ squared and grows with the mean. The dimensionless measure used throughout is the squared coefficient of variation, also called the noise strength: $$ \eta^2 = \frac{\text{Var}(P)}{\langle P \rangle^2} $$ so that $\eta$ is the standard deviation as a fraction of the mean. Copy numbers in eukaryotic cells span orders of magnitude — a transcription factor may sit at $\langle P \rangle \approx 10^{2}$ copies per cell while an abundant structural protein reaches $\langle P \rangle \approx 5\times 10^{5}$ — and measured noise strengths for most genes fall between about $0.05$ and $1$. One baseline anchors all of these numbers. If molecules are made and destroyed in independent single events, the copy number is Poisson-distributed, its variance equals its mean, and $$ \eta^2 = \frac{1}{\langle P \rangle} \qquad \text{(independent birth and death)} $$ Noise then falls as abundance rises: $10^{2}$ copies gives $\eta = 0.1$, and $10^{4}$ copies gives $\eta = 0.01$. Real genes are noisier than this, often far noisier at low expression, so the copy number is super-Poissonian, $\text{Var}(P) \gt \langle P \rangle$. The commonest reason is transcriptional bursting: a promoter is not continuously available, so transcripts arrive in intermittent groups rather than one at a time, and a single stochastic event at the promoter delivers several molecules at once. Arrivals in groups are more variable than arrivals one by one, which is enough to push the variance above the mean; the detailed kinetics of the switching are not needed here. Decomposing Noise: Intrinsic and Extrinsic Sources The total variability, $\eta_{\text{tot}}^2$, divides into two categories that differ in whether they act on one gene or on all of them. Intrinsic noise arises from the stochasticity of the reactions in a single gene's own pathway: the random timing of transcription initiation, of translation, and of the degradation of each species. Two identical genes in the same cell would still differ in output because of it, since their molecular events are separate draws. Extrinsic noise arises from fluctuations in the cellular context shared by both: the concentrations of RNA polymerase, ribosomes, ATP and amino acids, the activity of global regulators, and differences in cell size, growth rate and cell-cycle position. These act on many genes at once and therefore act on both copies together. If the two contributions are statistically independent, their variances add, and dividing by the squared mean gives the decomposition on which everything below rests: $$ \eta_{\text{tot}}^2 = \eta_{\text{int}}^2 + \eta_{\text{ext}}^2 $$ It is worth being exact about what the first term will turn out to measure. Operationally, $\eta_{\text{int}}^2$ is the variability that is uncorrelated between two copies of a gene, and $\eta_{\text{ext}}^2$ the part they share. The estimator does not know why two copies differ, only that they do, so anything independent between the two channels lands in $\eta_{\text{int}}^2$: detector and image-segmentation error, differences in the folding and maturation kinetics of the two fluorescent reporters, photobleaching, and any difference between the two chromosomal insertion sites. Elowitz and colleagues measured the instrumental contribution separately and subtracted it for precisely this reason. Read strictly, $\eta_{\text{int}}^2$ is an upper bound on gene-intrinsic stochasticity rather than a measurement of it. The