Chapter 08 of 08 · the confirmatory study

The Study

Everything The Journey built, assembled under a frozen contract: the world of 01, the POMDP of 02, the oracle of 03, the cloning of 04, the network of 05, the corruptions and recovery labels of 06, and the measurement discipline of 07.

The design

A base policy is trained on shared demonstrations, then cloned bit-exactly into three arms. The two full-budget arms each receive additional revealed oracle labels one as fresh demonstrations, one as recovery labels at learner-visited post-corruption states, with update counts, per-update target exposure, and replay rules matched exactly (the leak measured in chapter 06 is closed here by one-target-per-window training items). The whole pipeline is replicated as independent bundles.

The endpoint, frozen before any result

Success on eligible unseen scenarios (retained from candidates by a preregistered rule), each with one scheduled corruption from the unseen operator at held-out times, counted intention-to-treat, compared as per-bundle paired differences against a smallest effect of interest. What each evaluation slice can and cannot support:

The result

Success across all slices

Read the matrix with chapter 07's eyes: recovery training does not trade clean performance away; it leads everywhere, and loses almost nothing under the never-trained-on corruption. Precision: achieved half-width against the ±5 pp target; sensitivity interval (cluster bootstrap) .

Where recovery still loses

The same rule that picked the contrast on the front page picks this one: the first unseen scenario where the corruption lands and the recovery-trained policy misses it.

The first eligible unseen scenario where the recovery arm missed a delivered corruption.

Limitations, unchanged

In silico, symbolic, discrete, oracle-supervised. “Unseen” is the one other derangement of a three-action set. Findings on the matched slice are perturbation-family-specific. The treatment is DAgger-style recovery-state aggregation: masked behavioral cloning under budget accounting, not reinforcement learning, and no claim here transfers to physical robots without continuous control, imperfect perception, non-scripted supervision, and safety constraints this study deliberately excludes.

Three costs belong beside the benefit, and the report quantifies each. The budgets that were matched are label budgets: acquiring recovery labels took several times more oracle queries, because most recommendations produced during a learner rollout fall outside the reveal window. Among the episodes it solves, the recovery arm's paths are slightly longer relative to the oracle. And every failure in this environment is a step-limit timeout, because nothing here is irreversible, which is exactly the case where recovery labels should matter most.

The Journey, closed

Under equal additional label budgets and matched optimization, labeling the states a learner actually visits after an error beat collecting more expert demonstrations, by , 95% interval , on the frozen unseen endpoint. Every ingredient of that sentence is now yours to audit: the full report and the reproduction path are one click away.