Laya vs Jev

Jev 1.13.0 is a closed API. Laya is open weights you run yourself. The tables below are copied from published sources. This page does not re-run the benchmarks.

Read the 0.766 cell with the base-checkpoint row next to it. The README’s fine-tuned laya-typed-decisions scores 0.766 on 2,000 decisions. The English and multilingual bases score 0.362 and 0.352 on those same decisions, under a 0.461 majority-class baseline. pip install laya does not apply that fine-tune.

Official self-test

README section “Laya (with routing) vs Jev”. Laya figures are what the project says Router().predict returns. Jev figures are third-party published and were not measured in the Laya repository. Latency sources named in the README: AbdelStark/jev-benchmarks and nibzard/decision-model-benchmark.

The speed section of the same README also says a single question is roughly 6–7× faster against the 236–276 ms band. The comparison table itself prints 7.8×. Both sentences are in the README.
Jev 1.13.0Laya (routed)README note
typed-decisions, 2,000 decisions0.7270.766+0.039
AG News, 4 labels0.9100.950+0.040
DAIR Emotion, 6 labels0.4800.595+0.115
Banking77 (72 vs 77 labels)0.8700.425Jev leads on >20 options
ECE, lower better0.2460.0813× better (post-temperature)
p50 latency, 1 question236–276 ms32.8 ms7.8× faster
Languages usableno published benchmark45 of 51MASSIVE, above 3× random
Weightsclosed APIApache 2.0
Cost$0.042 / 1M tokens$0 self-hosted

Source: NandhaKishorM/laya README. On DAIR Emotion the README says Jev assigned zero probability to the true label on 16% of examples.

Official typed-decisions, all three checkpoints

400 cases, 2,000 decisions, four workflows. Same README section. Jev and the baselines are marked as published reference rows.

ModelAccuracySoft accBrierECEScore MAE
laya-typed-decisions0.7660.4710.0620.2130.242
laya0.3620.3320.3160.1750.694
laya-multilingual0.3520.3280.4630.3140.760
Jev 1.13.0 (published)0.7270.5800.1480.1440.391
Teacher self-agreement ceiling0.735
Per-question majority class0.461
Random guess0.318

The README says the fine-tuned checkpoint wins all four workflows (invoice processing 0.804, security incidents 0.766, customer service 0.764, agent-trace observability 0.730) and, by primitive, noul 0.857, choice 0.733, score 0.723. It also says the fine-tune clears the 0.735 teacher ceiling, so part of that gain can be fit to label noise. Full per-language detail is in BENCHMARKS.md.

Where the README says Jev leads

Why the router exists

Shared benchmark in the README: 17,416 questions, one T4, identical questions per model. Khmer on the English checkpoint is 0.000 accuracy at 0.952 confidence, which is why confidence gating cannot replace routing.

Tasklayalaya-multilingualRouter
MASSIVE intent, English0.7830.6570.783
MASSIVE intent, 13 other languages0.3060.4510.451
XNLI, English0.8600.8430.860
XNLI, 14 other languages0.5210.7310.731
Languages usable (>3× random)23 / 5145 / 5145 / 51
Latency, 1 question (T4)39.5 ms32.8 ms32.8 ms
Latency, 10 questions batched158.6 ms72.3 ms72.3 ms

Third-party: Luni/laya-jev-benchmark

Hugging Face dataset card Luni/laya-jev-benchmark. The card says Laya was remeasured on one RTX 5090. Rows marked published are quotes, not a Jev API rerun by this site or, on the card’s own account, by that author.

PhishNChips, 2,000 emails

AreLit/PhishNChips as reported on the Luni card. 1,000 phishing and 1,000 legitimate.
ModelAccuracyECEAUROCRecallp50
Laya, raw0.5050.4410.6780.0129 ms
Laya, Platt-calibrated0.611 0.679 9 ms
Jev (published)0.6260.1540.6890.432239 ms
Claude Haiku 4.5 (published)0.8130.0970.8370.764687 ms

typed-decisions, 400 cases

The card’s remeasure. 0.360 with no fine-tune is close to the README’s 0.362 English base, not a second official number.
ModelAccuracyECEms per case
Laya, no fine-tuning0.3600.17515.9
Laya fine-tuned on this task0.7670.21216.4
Jev 1.13.0 (published)0.7270.144710
Teacher self-agreement0.735

The same card reports RTX 5090 fp16 latency after warmup of 10.7 ms p50 for one question, 42.6 ms for 10, and a 14.9 s cold load. That hardware is not the T4 used in the README.

Third-party fine-tune: Cahol/laya-banking77-v1

This is not a Laya vs Jev run. The card evaluates a LoRA fine-tune on the official 3,080-example BANKING77 test, all 77 labels, and says the result does not show that Laya is better than Jev or a general LLM.

Source: Cahol/laya-banking77-v1. Temperature was fit on a held-out calibration split and does not change top-1.
ModelAccuracyMacro F1Top-3 accuracy
Laya English base, same 77-label protocol45.91%42.90%69.42%
Cahol fine-tune85.55%85.53%96.43%

Sources

Try the labeled demo in the playground, or install the package from Get started.

Laya vs Jev · Laya AI