Reconceptualising Entropy in Statistical Physics — Epoche C1
The habit of glossing entropy as 'disorder' survives because it gives the right answer for the cases textbooks introduce first — a gas expanding into a vacuum, a crystal melting — and it fails badly for cases that are not much more exotic. A suspension of hard spheres crystallises as it is compressed, gaining order and entropy simultaneously; ordinary ice retains a large entropy as its temperature is driven towards absolute zero, although it is a crystal; and the entropy change on mixing two gases depends on whether the experimenter can tell them apart. This essay defends the definition of entropy as the logarithm of the number of accessible microstates, shows what that definition delivers in each of those cases, and corrects one claim in the earlier version of this argument: the low-temperature ferromagnet, offered there as an example of an ordered state with high entropy, is not one. What the Formula Counts, and Why It Takes a Logarithm Everything below depends on three definitions, so they are given first and in full. A microstate is a complete specification of the system at the microscopic level — in classical mechanics, the position and momentum of every particle; in quantum mechanics, a particular state vector from an energy eigenbasis. A macrostate is a specification in terms of the few variables an experimenter actually fixes or measures: energy, volume, particle number, perhaps magnetisation. Accessible microstates are those compatible with the macrostate — those that satisfy the imposed constraints. Boltzmann's relation then defines $$S = k_B \ln \Omega,$$ where $\Omega$ is the number of accessible microstates and $k_B$ is the Boltzmann constant, fixed since the 2019 revision of the SI at exactly $1.380649 \times 10^{-23}\,\mathrm{J\,K^{-1}}$. Two features of that expression are not arbitrary. The logarithm is forced by additivity. If two independent systems are placed side by side, every microstate of the first can occur with every microstate of the second, so the combined count is the product $\Omega_{12} = \Omega_1 \Omega_2$. Thermodynamic entropy, however, is extensive: two identical bricks have twice the entropy of one. The only function turning products into sums is the logarithm, and $\ln(\Omega_1\Omega_2) = \ln\Omega_1 + \ln\Omega_2$ delivers exactly that. The constant $k_B$ is then a unit conversion and nothing more: $\ln\Omega$ is a pure number, whereas the thermodynamic entropy defined by Clausius through $dS = \delta Q_{\mathrm{rev}}/T$ carries units of energy divided by temperature, and $k_B$ supplies them (Callen 1985). Note what the definition does not mention: any notion of pattern, symmetry, or how a configuration looks. It counts, and whether the counted configurations look tidy is irrelevant to the count. Gibbs, Shannon, and Why the Two Entropies Must Be Kept Apart The earlier version of this essay said that the 'disorder' gloss blurs the distinction between information-theoretic and statistical-mechanical entropy. That is right, but the reason needs stating precisely, because the two are not different formulae. The Gibbs entropy of a probability distribution $\{p_i\}$ over microstates is $$S_G = -k_B \sum_i p_i \ln p_i,$$ which is Shannon's information entropy with $k_B$ in place of a choice of logarithm base. When the distribution is uniform over $\Omega$ accessible microstates, so that $p_i = 1/\Omega$ for each, the sum has $\Omega$ identical terms and $$S_G = -k_B \sum_{i=1}^{\Omega} \frac{1}{\Omega}\ln\frac{1}{\Omega} = k_B \ln \Omega,$$ recovering Boltzmann's expression exactly. The formulae agree; what differs is their argument. Boltzmann's $S$ is a function of a macrostate — of a region of phase space carved out by the constraints. Gibbs's $S_G$ is a functional of a distribution . The difference is not pedantic, and Liouville's theorem shows why. That theorem states that under Hamiltonian evolution the phase-space density $\rho$ is constant along any trajectory, so phase-space volume is neither created nor destroyed. It follows that $S_G = -k_B\int \rho \ln \rho \, d\Gamma$, computed from the exact fine-grained density, is rigorously constant in time. A quantity that cannot change cannot be the quantity that increases in the second law. The increase must therefore come from coarse-graining: from asking how much phase-space volume is compatible with the macroscopic description, which is precisely what $\Omega$ measures. Reading 'entropy' as an undifferentiated measure of disorder makes this distinction invisible, and with it the reason the second law needs a macrostate at all (Pathria and Beale 2011). Order Without Energy: Hard Spheres Crystallise to Gain Entropy The cleanest refutation of 'order equals low entropy' comes from a system chosen so that energy cannot be responsible for anything. Take $N$ impenetrable spheres in a box: the interaction energy is zero for every non-overlapping configuration and infinite for any overlap. Every allowed configuration therefore has the same energy, and the Helmholtz free energy $F = U - TS$ reduces to $F = -TS$ at fixed $U$. Since equilibrium minimises $F$, the equilibrium phase is simply whichever phase has the most configurations. Entropy is the only thing deciding. Alder and Wainwright (1957) integrated the equations of motion for a few hundred such spheres and found that above a packing fraction of roughly $0.5$ the fluid spontaneously develops crystalline order. The maximum packing fraction for spheres on a face-centred cubic lattice is $\pi/\sqrt{18} \approx 0.74$, so the transition occurs at about two-thirds of close packing — crowded, but far from jammed. This was a computational experiment: what was measured was the pressure as a function of density, and what was found was a van der Waals-type loop signalling coexistence between a fluid and an ordered solid. The mechanism is countable. In a dense disordered fluid each sphere is boxed in by neighbours at haphazard distances, so the free volume it can exp