Chapter 07 of 08 · teaching experiments

Believing a number

The question is sharp by now. What remains is everything that makes an answer trustworthy, and every piece of that discipline exists because a specific, demonstrable bias would otherwise creep in. This chapter demonstrates the biases.

Intention-to-treat

Analyze by what was scheduled, not by what happened. Every perturbed evaluation episode has a scheduled corruption time; sometimes the episode ends first and the corruption is never delivered. Delivery depends on the policy's own behavior, a post-treatment variable, so conditioning on it compares filtered, non-comparable subsets. In a simulation built so the true effect is by construction, the ITT estimator's bias is while the delivered-only estimator is off by , a bias the size of the effects this field measures, produced by nothing but an innocent-looking filter.

Two histograms of simulated estimates: the ITT distribution centered on the true effect, the per-protocol distribution shifted away from it.
2,000 simulated replicates against a known truth (vertical line).

Matched budgets

The study's currency is revealed oracle labels, and the match was deliberately broken to see what happens: rerunning chapter 06's recovery arm with labels (double the budget) lands at against matched. The feared inflation did not materialize; the base sits near its headroom ceiling, so doubled supervision bought nothing measurable in one replicate. That is the deeper point: an unmatched design attributes to the method whatever the extra budget did or did not do, and a small unfrozen run cannot even tell you which way the confound cuts. Matching is a design necessity for attribution, not an empirical convenience.

Three bars: the matched extra and recovery arms, and the double-budget recovery arm.
The unmatched design, run for comparison.

Paired replicates

Every pipeline replicate re-rolls collection and training, shifting both arms together; analyzing per-replicate differences cancels that shared noise. In simulation at six replicates, the paired analysis detects a five-point gap with power against unpaired. The entire reason the study runs six bundles and reports a paired interval.

Histograms of interval widths: paired intervals are much narrower than unpaired ones.
Interval widths over 2,000 simulated replicates.

Frozen means mechanical

A frozen protocol is a set of mechanisms, not a promise: unresolved design choices block confirmatory runs; the whole design hashes to one canonical fingerprint; data files chain their hashes so a quiet edit is caught at its exact row; the test set opens once, against a receipt; and the interpretation of the final interval was written down before the interval existed. Try the fingerprint yourself:

Summary

Four biases with names (selective delivery, unequal budgets, shared-noise dilution, post-hoc interpretation) and the mechanism that blocks each one. You can now audit the study's every sentence.