Ribosome Competition and Non-Proportional Gene Expression — Epoche C1
Doubling the amount of a messenger RNA in a cell does not double the amount of the protein it encodes, and the size of the shortfall can be calculated rather than merely asserted. If the gene in question already commands a fraction $\phi$ of the cell's translational capacity, then doubling its transcript raises its protein output by a factor of $2/(1+\phi)$, not 2 — a result derived below from a three-line model. For a gene taking one tenth of the capacity that is a 1.82-fold rise; for a gene taking half of it, only 1.33-fold. The reason is that ribosomes, the molecular machines that read messenger RNA and polymerise protein, exist in a fixed number, and every transcript is drawing on the same pool. This essay works out how tight that constraint is, what the competition model predicts, and how the predictions stand against measurement — including one measurement that appears at first to contradict them. How tight the ribosome budget is Before modelling competition it is worth establishing that ribosomes are genuinely scarce, since the argument collapses if they are in excess. Two calculations settle it, and both use only quantities collected in Milo and Phillips's Cell Biology by the Numbers (2015). The first is a headcount. A fast-growing Escherichia coli cell contains about $3\times10^{6}$ protein molecules of average length $300$ amino acids, and doubles in about $1200\,\mathrm{s}$. A bacterial ribosome elongates a polypeptide at roughly $20\,\mathrm{aa\,s^{-1}}$. The cell must therefore polymerise $3\times10^{6}\times300 = 9\times10^{8}$ amino acids in one doubling, while one ribosome supplies $1200\times20 = 2.4\times10^{4}$ of them, requiring $$N_R = \frac{9\times10^{8}}{2.4\times10^{4}} \approx 3.7\times10^{4}$$ ribosomes. Each factor is doing identifiable work: the numerator is the total peptide bond count the cell must form, the denominator is one machine's output over the same interval, and the answer agrees with the measured ribosome complement of fast-growing E. coli , which is in the tens of thousands. There is no slack in this figure: it is what the cell needs, not what it happens to have. The second calculation shows why that must be so. Let $\phi_R$ be the fraction of the cell's protein mass that is ribosomal protein, and $\lambda$ the growth rate. Since protein is made only by ribosomes, the rate of protein mass accumulation is the number of ribosomes times their output, and dividing by the total protein mass gives $$\lambda = \kappa\,\phi_R,$$ where $\kappa$ is a constant with units of inverse time — the maximum protein made per unit of ribosomal protein per unit time. Ribosomes are the only means of making more ribosomes, so the fraction of the proteome devoted to them must rise in strict proportion to how fast the cell grows. Scott and colleagues (2010) confirmed this linear relation experimentally across growth conditions in E. coli , and it implies a second law that matters here directly: forcing the cell to express a protein that does nothing useful, at mass fraction $\phi_U$, leaves less of the budget for $\phi_R$ and therefore reduces $\lambda$ in proportion. The cost of expression is not metaphorical. It is a subtraction from a conserved quantity, and it is measurable as lost growth. Ceroni and colleagues (2015) built a fluorescent capacity monitor that reports this burden directly, allowing synthetic constructs to be ranked by how much of the host's translational capacity they consume. The competition model, derived The earlier version of this essay proposed a formula for the translation rate of transcript $i$ and described it as qualitative. It can be derived, and deriving it corrects two things in the version given. Let $R_T$ be the total ribosome concentration, $R_f$ the free (unbound) portion, and $m_j$ the concentration of transcript species $j$, which binds free ribosomes at its initiation site with dissociation constant $K_j$ — a small $K_j$ meaning a strong initiation site that captures ribosomes readily. At equilibrium the fraction of species $j$ carrying an initiating ribosome is $(R_f/K_j)/(1+R_f/K_j)$. Ribosomes are conserved, so $$R_T = R_f + \sum_j m_j\,\frac{R_f/K_j}{1+R_f/K_j}.$$ Take the regime in which initiation, not elongation, limits output, so that $R_f$ is small compared with each $K_j$ and the occupancy fractions reduce to $R_f/K_j$. Then $R_T \approx R_f\big(1+\sum_j m_j/K_j\big)$, and the protein output of species $i$, which is its occupancy times its elongation rate constant $k_{el,i}$, becomes $$J_i \;=\; \frac{k_{el,i}\,(m_i/K_i)\,R_T}{1+\sum_{j} m_j/K_j}.$$ This has the shape the earlier version gave, with the competition coefficients identified as $1/K_j$ — they are not free parameters but the reciprocal binding strengths of the competing initiation sites. Two corrections follow. First, the total ribosome pool $R_T$ appears explicitly in the numerator; the earlier expression mentioned the pool in the surrounding text but omitted it from the formula, which then had no way of representing a change in ribosome number. Second, the sum in the denominator runs over all species including $i$ itself, not over $j \neq i$. A transcript competes with its own copies, and this is not a technicality: it is precisely the term that makes the response to overexpression sub-proportional. What the model predicts Write $a = m_i/K_i$ for the demand contributed by species $i$ and $D = \sum_j m_j/K_j$ for total demand, so that $J_i \propto a/(1+D)$. Define $\phi = a/(1+D)$, the share of ribosome-binding demand attributable to species $i$. Doubling $m_i$ doubles $a$ and increases $D$ by $a$, so the new output relative to the old is $$\frac{J_i(2m_i)}{J_i(m_i)} = \frac{2a}{1+2a+B}\Big/\frac{a}{1+a+B} = \frac{2(1+a+B)}{1+2a+B} = \frac{2}{1+\phi},$$ with $B$ the demand from all other species. The limiting cases are the check on the algebra. A transcript that is a negligible part of the load has $\phi \to 0$ and doubles its protein, recovering naive proporti