Task metric on the full test split; paired on identical examples. flips = examples lost/gained vs baseline.
TinyZero split · equation-correctness. Zero losses on every mask condition; mask_bon beats free BoN with a validity guarantee.
IBM scorer. Full Ctrl-G (λ=0.1) ties mask-only FSM with perfect validity and the best Partial of any condition.
FSM approaches mask-only from below as λ→0; validity is perfect at λ≥0.5. Dashed line = mask-only reference.
Chain-validity is the metric masking cannot move (its residual is truncation); the guide converts it.
Distance from the true constrained distribution vs guide capacity. Guidance beats masking only past a quality bar; oracle = exactly 0.
Upper bound on mask-addressable mass: walker-invalid (parse) + incomplete vs semantic failures, per unconstrained run.
Local 9B/8B + constraint vs frontier API models. Bars: frontier range per task; dots: our best constrained local result.
~/.local; fixed via shared PLTADDONDIR. Labels rescored; Llama baseline 0.006→0.193 after restatement-aware verification.