Ctrl-G × NWA — benchmark suite & HMM program

Headline — constrained decoding vs own baseline, per benchmark

Task metric on the full test split; paired on identical examples. flips = examples lost/gained vs baseline.

table view

Countdown (Qwen3.5-9B, n=256) — all conditions

TinyZero split · equation-correctness. Zero losses on every mask condition; mask_bon beats free BoN with a validity guarantee.

table view

NESTFUL (Llama-3.1-8B, test-200) — the HMM flagship

IBM scorer. Full Ctrl-G (λ=0.1) ties mask-only FSM with perfect validity and the best Partial of any condition.

table view

NESTFUL guidance-strength sweep (train-100)

FSM approaches mask-only from below as λ→0; validity is perfect at λ≥0.5. Dashed line = mask-only reference.

table view

nested_arith (Qwen3.5-9B, n=384) — termination steering

Chain-validity is the metric masking cannot move (its residual is truncation); the guide converts it.

table view

Exact-KL faithfulness (bounded Dyck-2, |L|=1347)

Distance from the true constrained distribution vs guide capacity. Guidance beats masking only past a quality bar; oracle = exactly 0.

table view

Structural share of baseline failures (diagnostic)

Upper bound on mask-addressable mass: walker-invalid (parse) + incomplete vs semantic failures, per unconstrained run.

table view

Frontier context (provider defaults, same files & verifiers)

Local 9B/8B + constraint vs frontier API models. Bars: frontier range per task; dots: our best constrained local result.

table view

What broke and how it was fixed