Fine-tune before you quote 0.766
The package runs zero-shot. The number people repeat is a checkpoint trained on that benchmark’s own split. This page is the README’s fine-tune note, not a new training recipe.
What install gives you
On typed-decisions (2,000 decisions, four workflows) the README and the Hugging Face model card report 0.766 for laya-typed-decisions. The English base is 0.362 and the multilingual base is 0.352. A per-question majority class scores 0.461, and a random guess scores 0.318. The card’s own line is that the base checkpoints sit below the majority baseline, and that the capability on this benchmark comes from fine-tuning. It is not a zero-shot result.
typed-decisions is not chosen by the router unless you pass model="typed-decisions" or turn on automatic task detection. It is fit to four synthetic workflows. Using it as a silent default for your tickets is the wrong checkpoint.
RLCD in the notebook the project ships
laya_finetune_typed_decisions_2xT4_kaggle.ipynb is the loop the README and the model card describe: build the dataset, train with RLCD (proper-scoring-rule rewards, GRPO-style policy gradient), fit calibration temperatures, evaluate, and push to the Hub. It targets Kaggle’s free 2× T4 GPUs. The README says runtime is roughly 4–5 hours for 4 epochs over about 30k questions. This page does not add hyperparameters beyond that description.
- The notebook enables gradient checkpointing on the encoder and the decision head.
- It fits one temperature per type (
choice,score,noul) and removes inheritedtemperature_by_options. Leaving the old bucket values in place makes them win at inference and hide the new fit. - Calibration samples in that notebook come from its training items. The README says to evaluate on separate held-out data before claiming an improvement. Already published checkpoints are not rewritten.
Both shipped checkpoints are over-confident. Refitting one temperature per question type and option count moves mean ECE from 0.466 to 0.081 on laya and from 0.314 to 0.106 on laya-multilingual. The multilingual checkpoint ships with no fitted temperatures.
Use the question shape first
Label rows in the same three types the model answers. A support router is a choice, a score, and a noul, which is what the support-triage workbench edits. When the label set is large, read many labels before you spend the training run on a question the head budget cannot read.
The full scoreboard, including where the fine-tune still trails Jev on soft accuracy and raw ECE, is on Laya vs Jev.