Get started

Install the package, preload the checkpoints you will actually serve, and call Router().predict. The playground on this site does not load the weights.

Install

Python 3.10 or newer. The README names that floor because current huggingface_hub, transformers, and torch releases require it. Declared dependencies are torch, transformers, safetensors, huggingface_hub, and numpy. Optional extras include laya[serve], laya[mcp], laya[onnx], and laya[fast].

PyPI
python -m pip install laya
venv, macOS or Linux
python3 -m venv .venv
.venv/bin/python -m pip install laya
.venv/bin/python -I -c "import laya; print(laya.__version__)"
uv
uv venv --python 3.12
uv pip install laya

-I keeps a local source tree from masking a missing install. If you need a CPU-only or GPU-specific PyTorch wheel, install it in the same environment before laya, following the PyTorch guide. A missing rl_agent_config.json means the checkpoint directory is incomplete. That file ships with the weights. It is not something you create by hand.

Package page: pypi.org/project/laya. Development install: pip install "git+https://github.com/NandhaKishorM/laya.git".

Hardware

The README does not publish a minimum RAM or VRAM figure. Parameter counts are the size numbers it does publish. Latency below was measured on a Tesla T4 unless a row says CPU.

Checkpoint table from the README.
CheckpointEncoderParamsContextUse it for
layaModernBERT-large421M512English
laya-multilingualmmBERT-base322M1024, up to 8192100+ languages
laya-typed-decisionsModernBERT-large421M1024typed-decisions workflows
Speed on a Tesla T4, one call, from the README.
Questions per calllayalaya-multilingual
139.5 ms32.8 ms
584.5 ms40.1 ms
10158.6 ms (15.9 ms/q)72.3 ms (7.2 ms/q)
50771 ms337 ms (6.8 ms/q)

With Router(preload=True) the README reports 32.8 ms on GPU and 193–464 ms on CPU. A 4,000-token input is about 1.7 s on an Apple GPU. Long documents belong on model="multilingual" with max_len=8192. The checkpoint’s default limit is 1,024 and will cut them off.

Preload pitfalls

A cold checkpoint build costs seconds. Language detection costs microseconds. The lazy default keeps two checkpoints resident, english and multilingual, because those are the two automatic routing chooses between. max_loaded=1 rebuilds the checkpoint it just evicted on every switch: the README measures a 7.4 s median reload on CPU and 10.3 s on a T4.

ModePer-request latencyReloads
Router() lazy, max_loaded=2detection only (<1 ms) on a switch, after each language’s first load1 the first time a language appears
Router(max_loaded=1)7 to 10 s on every language switch1 per switch
Router(preload=True)32.8 ms GPU / 193–464 ms CPUnone
Preload
router = Router(preload=True)
router = Router(preload=True, device="cuda")
router.preload(["english", "multilingual"])
router.attach("english", existing_agent)  # avoid a second copy in VRAM
router = Router(max_loaded=1)             # reloads on every switch
router.unload()

Very short Latin text such as “Quero cancelar” often has no language signal and follows default, which is english unless you set Router(default="multilingual"). The English checkpoint collapses on scripts it cannot read. The README’s Khmer figure is 0.000 accuracy at 0.952 confidence, so a confidence gate will not catch that failure. typed-decisions is never chosen unless you pass model="typed-decisions" or turn on auto_task_detection.

A separate community card, Luni/laya-jev-benchmark, reports a 14.9 s cold load on an RTX 5090. That is their machine, not the T4 table.

Quickstart

Shortest path from the top of the README. Comments are the README’s, not a run from this site.

Short quickstart
from laya import Router

router = Router()  # downloads a checkpoint on first use; Router(preload=True) loads all three up front

state = "Hi, we were billed twice for March. Please refund the duplicate today or we will cancel our plan."
questions = {
    "department": {"type": "choice", "instructions": "Which department should handle this?",
                   "criteria": {"billing": "invoices, payments, refunds",
                                "technical": "bugs, outages, system errors",
                                "other": "everything else"}},
    "urgency": {"type": "score", "instructions": "How urgent is this?",
                "criteria": ["not urgent", "soon", "blocking"]},
    "churn_risk": {"type": "noul", "instructions": "Does the user threaten to cancel or leave?"},
}

result = router.predict(state, questions)
print(result["answers"]["department"]["choice"])  # billing
print(result["answers"]["churn_risk"]["noul"])    # probability the answer is yes
print(result["routing"]["model"])                 # english

Route mode, the recommended production entry point:

Router quickstart
from laya import Router

# Preload checkpoints into memory for instant sub-35ms routing
router = Router(preload=True)

# 1. State in any language or schema
state = {
    "from": "user@acme.com",
    "subject": "Duplicate charge on invoice #4411",
    "body": "Hi, we were billed twice for March. Please refund the duplicate today or we will cancel our plan."
}

# 2. Define your typed questions
questions = {
    "department": {
        "type": "choice",
        "instructions": "Which department should handle this request?",
        "criteria": {
            "billing": "invoices, payments, refunds",
            "technical": "bugs, outages, system errors",
            "sales": "pricing, new contracts",
            "other": "everything else"
        }
    },
    "urgency": {
        "type": "score",
        "instructions": "How urgent is this request?",
        "criteria": ["not urgent", "soon", "critical deadline or blocking issue"]
    },
    "churn_risk": {
        "type": "noul",
        "instructions": "Does the user threaten to cancel or leave?"
    },
    "refund_requested": {
        "type": "noul",
        "instructions": "Does the user explicitly request a refund?"
    }
}

# 3. English state -> automatically routed to laya (ModernBERT-large, 39.5 ms)
res_en = router.predict(state, questions)
print("Department :", res_en["answers"]["department"]["choice"])  # -> billing (confidence: 0.94)
print("Routing    :", res_en["routing"]["model"])                 # -> english

# 4. Hindi state -> automatically routed to laya-multilingual (mmBERT-base, 32.8 ms)
res_hi = router.predict({"body": "मुझसे दो बार शुल्क लिया गया, कृपया पैसे वापस करें।"}, questions)
print("Department :", res_hi["answers"]["department"]["choice"])  # -> billing (confidence: 0.86)
print("Routing    :", res_hi["routing"]["model"])                 # -> multilingual

# 5. Explicit override when you want a specific checkpoint
res_td = router.predict(state, questions, model="typed-decisions")

Direct load, when the pipeline should stay on one checkpoint: laya.load("convaiinnovations/laya"), with subfolder="multilingual" or subfolder="typed-decisions". The README’s annotated sample for that path prints department billing (confidence 0.94), urgency 1.84 / 2.0, and churn noul 0.892. Treat those as sample comments. Real probabilities move with checkpoint, device, and wording.

Empty questions return answers: {} and zero token usage, and the state is not tokenized. confidence on choice and score is 1 minus normalised entropy. answer_confidence is the probability of the reported answer and is the number to gate on if you want one scale across question types. A threshold copied from Jev does not transfer. Fit temperatures on your own held-out data before you trust a cutoff. Both shipped checkpoints are over-confident.

Command line

laya CLI
laya "I was charged twice, please refund"            # route only, no download
laya "Refactor this service" --predict               # full answers
laya "Mein Konto wurde zweimal belastet" --lang de
laya "My payment failed twice" --preset triage
laya --batch tickets.txt --predict

Routing alone does not download a checkpoint. --predict does, and needs Hugging Face network access the first time. Presets shipped in the package are triage, email, guard, moderation, and router, via laya.triage_questions() and the matching helpers.

Limits worth knowing before you ship

Numbers and the wording of each limit are in the README honest-limits section. The sourced Jev comparison is on Laya vs Jev.

HTTP server

laya-serve speaks POST /v1/systemone, the same wire protocol the README attributes to TypeSafe’s Jev API. Three differences the README calls out: a shared option token budget rather than 255 options, every score level needs a description, and confidence is not Jev’s formula.

laya-serve
pip install "laya[serve]"
LAYA_DEVICE=cuda LAYA_PRELOAD=1 laya-serve

Also see the upstream docs, the Colab notebook, and the official Space if you want their weights in a browser.

Run locally · Laya AI