The Image Carried by Photons the Camera Never Sees — Epoche C2
Problem Mid-infrared imaging is where molecular fingerprints live: fundamental vibrational bands of C–H, O–H and C=O stretches all fall between roughly 3 and 10 $\mu$m. The instrumentation there is poor by comparison with the visible. The reason is dark current in the detector. In the diffusion-limited regime it goes as the square of the intrinsic carrier density, and therefore as $\exp(-E_g/k_BT)$ in the semiconductor band gap $E_g$. Silicon has $E_g = 1.12$ eV, and at room temperature $k_BT = 0.0259$ eV, so the factor is $\exp(-43.3) = 1.6\times10^{-19}$. Indium antimonide, one of the standard mid-infrared materials, has $E_g = 0.17$ eV at room temperature, widening to about 0.23 eV at 77 K, where $k_BT = 0.00664$ eV; its factor is $\exp(-34.6) = 9\times10^{-16}$. The ratio is $9\times10^{-16}/1.6\times10^{-19} = 5.6\times10^{3}$ — nearly four orders of magnitude worse than uncooled silicon, and that is after paying for liquid nitrogen. The obvious question is whether one can probe at 3.7 $\mu$m and detect in the visible. The obvious answer is that one cannot, because an image is formed by light that has interacted with the object and then reached the detector. This note records why that answer is wrong, what replaces it, and what the substitution actually buys. Units are SI, with wavelengths in nm and $\mu$m as in the source papers. Configuration The arrangement is the induced-coherence interferometer introduced by Zou, Wang and Mandel in 1991, used as an imager by Lemos and colleagues in 2014. A single pump laser is split and sent through two nonlinear crystals, NL1 and NL2, one after the other along separate arms. Each crystal can convert one pump photon into a pair by spontaneous parametric down-conversion — a signal photon and an idler photon, at wavelengths fixed by energy conservation, $1/\lambda_p = 1/\lambda_s + 1/\lambda_i$. Lemos and colleagues used $\lambda_p = 532$ nm and $\lambda_s = 810$ nm; the idler then follows as $1/\lambda_i = 1/532 - 1/810 = (810-532)/(532\times810) = 6.451\times10^{-4}$ nm$^{-1}$, giving $\lambda_i = 1550$ nm. The pump power is kept low enough that the probability of a pair from either crystal in one coherence time is far below one. Almost every detected event therefore involves at most one pair. The idler beam emerging from NL1 is aligned so that it passes through NL2 collinearly with the idler mode NL2 itself emits. This is the critical step: after NL2, nothing about the idler beam records which crystal made the pair. The two signal beams are recombined on a beamsplitter and imaged onto a camera. The idler is discarded — blocked, or simply never detected. The object is placed in the idler path between the two crystals. The relation that does the work Because the pair can come from either crystal and the two possibilities are indistinguishable at the idler, their amplitudes add. Write the object's complex amplitude transmission as $T$, so that $|T|^2$ is the intensity transmittance and $\arg T$ the phase it imposes. Tracing the idler out of the two-crystal state leaves a signal density matrix whose off-diagonal element between the two paths is precisely the overlap of the two idler modes — and that overlap is $T$. The detected signal rate is therefore $$R_s \;\propto\; 1 + |T|\,\cos\!\left(\varphi_s + \varphi_i - \varphi_p + \arg T\right),$$ where $\varphi_s$, $\varphi_i$ and $\varphi_p$ are the propagation phases accumulated in the signal, idler and pump arms between the crystals. Two features deserve emphasis. First, the fringe visibility equals $|T|$: the object's amplitude transmission is read directly off the contrast of an interference pattern in a beam that never met it. Second, the fringe position gives $\arg T$, so the same measurement yields a quantitative phase image — the refractive index map at the idler wavelength. The limiting cases confirm the reading. If the object is removed, $T = 1$ and the visibility is unity. If it is opaque, $T = 0$: the idler now carries which-crystal information, the two paths become distinguishable, and the fringes vanish entirely, leaving a flat sum of intensities. Nothing here requires the signal photon to have been anywhere near the sample. What is destroyed by the object is not a signal but an indistinguishability. It is worth being precise about the sensitivity, because it is easy to overclaim. Since $V = |T| = \sqrt{|T|^2}$, an object transmitting 90% of the intensity gives $V = \sqrt{0.9} = 0.949$, a visibility reduction of 5.1% for a 10% absorption. The method is therefore about half as responsive to weak absorption as a direct intensity measurement of the same beam would be. Its advantage is not sensitivity per photon. What is bought The advantage is that $\lambda_s$ and $\lambda_i$ are decoupled, tied only by the pump. Choose $\lambda_p = 532$ nm and put the probe in the mid-infrared at $\lambda_i = 3.7\ \mu$m; the signal follows from the same energy-conservation relation, $1/\lambda_s = 1/532 - 1/3700 = 1.6094\times10^{-3}$ nm$^{-1}$, so $\lambda_s = 621$ nm. That is squarely in the response band of an ordinary silicon camera, and the four-order-of-magnitude detector penalty computed above disappears. Kviatkovsky and colleagues demonstrated exactly this in 2020, and Kalashnikov and colleagues had earlier used the same principle for infrared spectroscopy read out with visible light. A second, less advertised benefit is the exposure delivered to the sample. At a plausible idler rate of $10^{7}$ s$^{-1}$ and a photon energy $hc/\lambda_i = 1239.8\ \mathrm{eV\,nm}/3700\ \mathrm{nm} = 0.335$ eV $= 5.37\times10^{-20}$ J, the optical power on the object is $10^{7}\times5.37\times10^{-20} = 5.4\times10^{-13}$ W, about half a picowatt. For photosensitive or biological samples this is a genuine argument, independent of any detector consideration. Limits observed Visibility is not the object alone. In practice $V = |T|\times\eta_{mode}\times\eta_{bal}$, where $\eta_{mode}$ is the spatial and spectral over