Chapter 07 of 08 · teaching experiments
Believing a number
The question is sharp by now. What remains is everything that makes an answer trustworthy, and every piece of that discipline exists because a specific, demonstrable bias would otherwise creep in. This chapter demonstrates the biases.
Intention-to-treat
Analyze by what was scheduled, not by what happened. Every perturbed evaluation episode has a scheduled corruption time; sometimes the episode ends first and the corruption is never delivered. Delivery depends on the policy's own behavior, a post-treatment variable, so conditioning on it compares filtered, non-comparable subsets. In a simulation built so the true effect is by construction, the ITT estimator's bias is while the delivered-only estimator is off by , a bias the size of the effects this field measures, produced by nothing but an innocent-looking filter.
Matched budgets
The study's currency is revealed oracle labels, and the match was deliberately broken to see what happens: rerunning chapter 06's recovery arm with labels (double the budget) lands at against matched. The feared inflation did not materialize; the base sits near its headroom ceiling, so doubled supervision bought nothing measurable in one replicate. That is the deeper point: an unmatched design attributes to the method whatever the extra budget did or did not do, and a small unfrozen run cannot even tell you which way the confound cuts. Matching is a design necessity for attribution, not an empirical convenience.
Paired replicates
Every pipeline replicate re-rolls collection and training, shifting both arms together; analyzing per-replicate differences cancels that shared noise. In simulation at six replicates, the paired analysis detects a five-point gap with power against unpaired. The entire reason the study runs six bundles and reports a paired interval.
Frozen means mechanical
A frozen protocol is a set of mechanisms, not a promise: unresolved design choices block confirmatory runs; the whole design hashes to one canonical fingerprint; data files chain their hashes so a quiet edit is caught at its exact row; the test set opens once, against a receipt; and the interpretation of the final interval was written down before the interval existed. Try the fingerprint yourself:
Summary
Four biases with names (selective delivery, unequal budgets, shared-noise dilution, post-hoc interpretation) and the mechanism that blocks each one. You can now audit the study's every sentence.