Laya vs Jev
Jev 1.13.0 is a closed API. Laya is open weights you run yourself. The tables below are copied from published sources. This page does not re-run the benchmarks.
Read the 0.766 cell with the base-checkpoint row next to it. The README’s fine-tuned laya-typed-decisions scores 0.766 on 2,000 decisions. The English and multilingual bases score 0.362 and 0.352 on those same decisions, under a 0.461 majority-class baseline. pip install laya does not apply that fine-tune.
Official self-test
README section “Laya (with routing) vs Jev”. Laya figures are what the project says Router().predict returns. Jev figures are third-party published and were not measured in the Laya repository. Latency sources named in the README: AbdelStark/jev-benchmarks and nibzard/decision-model-benchmark.
| Jev 1.13.0 | Laya (routed) | README note | |
|---|---|---|---|
| typed-decisions, 2,000 decisions | 0.727 | 0.766 | +0.039 |
| AG News, 4 labels | 0.910 | 0.950 | +0.040 |
| DAIR Emotion, 6 labels | 0.480 | 0.595 | +0.115 |
| Banking77 (72 vs 77 labels) | 0.870 | 0.425 | Jev leads on >20 options |
| ECE, lower better | 0.246 | 0.081 | 3× better (post-temperature) |
| p50 latency, 1 question | 236–276 ms | 32.8 ms | 7.8× faster |
| Languages usable | no published benchmark | 45 of 51 | MASSIVE, above 3× random |
| Weights | closed API | Apache 2.0 | |
| Cost | $0.042 / 1M tokens | $0 self-hosted |
Source: NandhaKishorM/laya README. On DAIR Emotion the README says Jev assigned zero probability to the true label on 16% of examples.
Official typed-decisions, all three checkpoints
400 cases, 2,000 decisions, four workflows. Same README section. Jev and the baselines are marked as published reference rows.
| Model | Accuracy | Soft acc | Brier | ECE | Score MAE |
|---|---|---|---|---|---|
| laya-typed-decisions | 0.766 | 0.471 | 0.062 | 0.213 | 0.242 |
| laya | 0.362 | 0.332 | 0.316 | 0.175 | 0.694 |
| laya-multilingual | 0.352 | 0.328 | 0.463 | 0.314 | 0.760 |
| Jev 1.13.0 (published) | 0.727 | 0.580 | 0.148 | 0.144 | 0.391 |
| Teacher self-agreement ceiling | 0.735 | ||||
| Per-question majority class | 0.461 | ||||
| Random guess | 0.318 |
The README says the fine-tuned checkpoint wins all four workflows (invoice processing 0.804, security incidents 0.766, customer service 0.764, agent-trace observability 0.730) and, by primitive, noul 0.857, choice 0.733, score 0.723. It also says the fine-tune clears the 0.735 teacher ceiling, so part of that gain can be fit to label noise. Full per-language detail is in BENCHMARKS.md.
Where the README says Jev leads
- High-cardinality labels. Banking77 is 0.870 for Jev on 72 labels and 0.425 for Laya on 77 labels at the default 256-token head budget, about 3 to 4 tokens per label. Jev allows up to 255 options. The README points at
predict_shortlistand a largerhead_max_lenas the Laya workarounds. - Soft distribution match on typed-decisions: Jev soft accuracy 0.580, Laya 0.471, even where Laya’s argmax accuracy is higher.
- Raw calibration. Before temperature scaling the base checkpoint’s ECE is 0.213 against Jev’s 0.144. The 0.081 ECE is after domain temperature fitting. Refitting one temperature per question type and option count moves mean ECE from 0.466 to 0.081 on
layaand from 0.314 to 0.106 onlaya-multilingual. The multilingual checkpoint ships with no fitted temperatures.
Why the router exists
Shared benchmark in the README: 17,416 questions, one T4, identical questions per model. Khmer on the English checkpoint is 0.000 accuracy at 0.952 confidence, which is why confidence gating cannot replace routing.
| Task | laya | laya-multilingual | Router |
|---|---|---|---|
| MASSIVE intent, English | 0.783 | 0.657 | 0.783 |
| MASSIVE intent, 13 other languages | 0.306 | 0.451 | 0.451 |
| XNLI, English | 0.860 | 0.843 | 0.860 |
| XNLI, 14 other languages | 0.521 | 0.731 | 0.731 |
| Languages usable (>3× random) | 23 / 51 | 45 / 51 | 45 / 51 |
| Latency, 1 question (T4) | 39.5 ms | 32.8 ms | 32.8 ms |
| Latency, 10 questions batched | 158.6 ms | 72.3 ms | 72.3 ms |
Third-party: Luni/laya-jev-benchmark
Hugging Face dataset card Luni/laya-jev-benchmark. The card says Laya was remeasured on one RTX 5090. Rows marked published are quotes, not a Jev API rerun by this site or, on the card’s own account, by that author.
PhishNChips, 2,000 emails
| Model | Accuracy | ECE | AUROC | Recall | p50 |
|---|---|---|---|---|---|
| Laya, raw | 0.505 | 0.441 | 0.678 | 0.012 | 9 ms |
| Laya, Platt-calibrated | 0.611 | 0.679 | 9 ms | ||
| Jev (published) | 0.626 | 0.154 | 0.689 | 0.432 | 239 ms |
| Claude Haiku 4.5 (published) | 0.813 | 0.097 | 0.837 | 0.764 | 687 ms |
typed-decisions, 400 cases
| Model | Accuracy | ECE | ms per case |
|---|---|---|---|
| Laya, no fine-tuning | 0.360 | 0.175 | 15.9 |
| Laya fine-tuned on this task | 0.767 | 0.212 | 16.4 |
| Jev 1.13.0 (published) | 0.727 | 0.144 | 710 |
| Teacher self-agreement | 0.735 |
The same card reports RTX 5090 fp16 latency after warmup of 10.7 ms p50 for one question, 42.6 ms for 10, and a 14.9 s cold load. That hardware is not the T4 used in the README.
Third-party fine-tune: Cahol/laya-banking77-v1
This is not a Laya vs Jev run. The card evaluates a LoRA fine-tune on the official 3,080-example BANKING77 test, all 77 labels, and says the result does not show that Laya is better than Jev or a general LLM.
| Model | Accuracy | Macro F1 | Top-3 accuracy |
|---|---|---|---|
| Laya English base, same 77-label protocol | 45.91% | 42.90% | 69.42% |
| Cahol fine-tune | 85.55% | 85.53% | 96.43% |
Sources
- Official README and BENCHMARKS.md
- Luni/laya-jev-benchmark
- Cahol/laya-banking77-v1
- AbdelStark/jev-benchmarks and nibzard/decision-model-benchmark, as cited by the README for Jev latency
Try the labeled demo in the playground, or install the package from Get started.