Many labels fail at the default budget
A choice question does not get one slot per label. The labels share a token budget. Past about 20 options, the README says Jev leads unless you change the question.
What Banking77 shows
In the README’s routed comparison, Jev 1.13.0 scores 0.870 on 72 Banking77 labels. Routed Laya scores 0.425 on 77 labels at the default head budget. The mechanism the README gives: options share head_max_len (192 tokens on English, 256 on the other two checkpoints). A 77-option question gets about (256 - 16) // 77, which is 3 to 4 tokens per label, so the texts become hard to tell apart.
Jev’s published cap is 255 options. laya-serve rejects more than 100 choice options before inference. That is a different limit from the token budget, and neither one makes 77 similar banking intents readable at the default width.
Three ways the README says to handle it
Widen one request
predict(..., head_max_len=512, max_len=1024) changes the budget for that call. The state then has fewer tokens left, because the window is split between options and the document. The CLI flags are --head-max-len and --max-len.
result = agent.predict(state, questions, head_max_len=512, max_len=1024)
Shortlist, then one forward pass
predict_shortlist keeps the top k labels with an embedding, then scores only those. embed_fn_from_agent mean-pools the encoder already loaded, so you do not download a second model. Probabilities are over the shortlist, not the original 77.
import laya
questions = {
"intent": {
"type": "choice",
"instructions": "Which banking intent is this?",
"criteria": {
"card_arrival": "where is my card",
"transfer_fee": "fee charged on a transfer",
# ...the rest of a large label set
},
}
}
result = laya.predict_shortlist(
agent,
{"text": "I was charged twice for a transfer"},
questions,
embed_fn=laya.embed_fn_from_agent(agent),
k=20,
)
result["shortlist"]["intent"]["labels"] # the top 20 labels sent to the model
Issue #102 reports that a top-20 zero-shot shortlist moved one BANKING77 run from 54.3% to 60.8% on the reporter’s setup. The README says those figures are the reporter’s and that the repository has not remeasured them.
Split the question
Ask a coarse choice first (card, transfer, account), then a fine choice on the branch you kept. Two small label sets stay inside the budget. The support template is this shape: four departments, not 77 intents.
A fine-tune is a different claim
Cahol/laya-banking77-v1 reports 85.55% accuracy, 85.53 macro F1, and 96.43% top-3 on the official 3,080-example test, against 45.91% for the English base on the same 77-label protocol. The card says this does not show that Laya is better than Jev or a general LLM. It is a specialised head, not the zero-shot router.
If the label set is your product, start with a shortlist or a split, measure on your data, then fine-tune. The rest of the Jev table is on Laya vs Jev.