Training curves and holdout metrics. All models share the Transformer-over-skater-tokens state encoder.
Distribution of action choices on the held-out game. The CQL policy is heavily conservative (rarely dumps, over-passes) because dumps are situationally optimal in ways the state encoder can't observe.