laya-multilingual (ONNX)

ONNX conversion of Convai Innovations' convaiinnovations/laya-multilingual, a non-autoregressive decision model (jhu-clsp/mmBERT-base encoder with a decision head). Give it a state (text or JSON) and typed questions (choice, score, noul); it returns a calibrated distribution per question in one forward pass. For the model itself, its training and its limits, see the original model card and the laya repository.

Files

File Size
onnx/model.onnx (+ _data), fp32 1.29 GB
onnx/model_fp16.onnx (+ _data), fp16 weights and activations, fp32 inputs/outputs 0.65 GB

One graph holds encoder and head (laya's own laya-ts export splits them in two):

  • inputs: input_ids [b, s] int64, attention_mask [b, s] int64, marker_pos [b, k] int64, marker_mask [b, k] bool, qtype [b] int64 (choice 0, score 1, noul 2)
  • outputs: logits [b, k] (one per [MASK] option marker, masked slots -1e4), act_logits [b, 2]

Each question is its own sequence ([CLS] {type} question: {question} [SEP] [MASK] opt … [SEP] state [SEP]). config.json carries the encoder config plus a laya section with everything a runtime needs: max_len (1024), head_max_len (256), the calibrated temperatures per question type and option count, and whether the tokenizer needs word-by-word encoding in Transformers.js (split_words).

Use it in the browser

With open-jev (Transformers.js, WebGPU):

import { OpenJev, choice, score, noul } from "open-jev";

const jev = await OpenJev.load({ model: "laya-multilingual" });
const answers = await jev.decide("Hi, we were billed twice for March. Please refund the duplicate.", {
  department: choice("Which department should handle this?", ["billing", "technical", "other"], {
    billing: "invoices, payments, refunds",
    technical: "bugs, outages, system errors",
  }),
  urgency: score("How urgent is this?", ["not urgent", "soon", "blocking"]),
  refund: noul("Does the user ask for a refund?"),
});

The open-jev laya family rebuilds laya's build_sequence (option texts, 48-token option cap, head budget) and applies the checkpoint's calibrated temperatures, as laya.Agent does.

Conversion and parity

torch.onnx.export (dynamo, opset 18) of DecisionModel.forward from laya.common, with dynamic batch, sequence and marker axes. Scripts are in conversion/.

  • ONNX Runtime (CPU) against PyTorch on real typed-decisions sequences, up to 2,048 tokens (past the 128-token sliding window): max |Δlogit| 3.1e-5.
  • open-jev in Node (onnxruntime-node) against laya's own Python Agent.predict on 15 cases (typed-decisions states; double spaces; the mask token inside text; over-long options; the head budget; Hindi, Chinese, Japanese and Spanish): token ids identical for every question; fp32 max |Δp| 1e-4; fp16 max |Δp| 0.035, no top answer changed.

Accuracy in the browser

typed-decisions test split (400 states, 2,000 questions), WebGPU on an M3 Pro, via open-jev:

dtype choice score noul overall median per state
fp16 29.7% 28.0% 49.8% 35.0% 0.51 s
fp32 29.5% 28.5% 49.8% 35.2% 0.67 s

fp16 is the default on WebGPU: it matches fp32 within 0.2 points and is faster. The base checkpoints are zero-shot here; Laya's card reports 0.362 for the English one.

License

Apache-2.0, as the original. Converted by @shreyask; all credit for the model to Convai Innovations.

Downloads last month
31
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for onnx-community/laya-multilingual-ONNX

Quantized
(27)
this model