Calibration

Two fields look like confidence. The README says they are not the same, and that Jev’s formula does not transfer. The playground prints both. Neither number on this site is a checkpoint probability.

Which number is the gate

The shipped checkpoints are over-confident

The README’s temperature refit, one temperature per question type and option count, moves mean ECE from 0.466 to 0.081 on laya and from 0.314 to 0.106 on laya-multilingual. The multilingual checkpoint ships with no fitted temperatures. Temperatures are clamped to [0.5, 5.0]. An invalid value falls back to 1.0.

The fine-tune notebook fits one temperature per type and removes inherited temperature_by_options, because the old buckets win at inference if you leave them. Those calibration samples come from the notebook’s training items. Evaluate on held-out data before you quote the fit. See fine-tune.

Do not gate on the act head

action.act_probability is not a usable signal. #185 and the README cite AUROC 0.30 for that head against AUROC 0.77 for confidence, on 396 labelled decisions. The demo sets act_probability to null so a copied JSON shape is not mistaken for that head.

Raw ECE on the typed-decisions self-test is another place the README is explicit: Laya’s fine-tuned checkpoint is 0.213 before temperature, Jev’s published figure is 0.144. Post-temperature, the routed comparison quotes ECE 0.081 against Jev’s 0.246. Those are different columns. The tables stay on Laya vs Jev.

Calibration · Laya AI