# Introducing VerQen, an open decision model

> Beyond the chat box: a small model that applies your rules in milliseconds, knows when to hand over, and is open down to its training data. With the full benchmark, including the tracks it loses.

By Ved Saini. Published 2026-09-27. Canonical: https://zorqelisai.com/blog/introducing-verqen-open-decision-model

![VerQen, beyond the chat box: a chat window dissolves into a prism that emits three typed decisions, a choice, a score and a yes or no, each with a probability.](https://zorqelisai.com/blog/verqen/verqen-launch.webp)

*State in, decision out. VerQen answers with one option and a calibrated probability, never with prose.*

Today we are releasing **VerQen**, an open decision model. It reads a situation, a question and a fixed list of options, and returns one option, a calibrated probability and a flag that says when it is not sure enough to act. It is 232M parameters, trained from scratch by ZorQelis AI, and it answers in 5 to 7 milliseconds on a consumer GPU. The weights, the 1.62M training decisions and the benchmark we measured it on are all public, under open licences.

## Key takeaways

- VerQen is an open decision model from ZorQelis AI: it reads a situation, a question and a fixed set of options, and returns one option with a calibrated probability and a flag for when to escalate. It never generates text.
- It has 232M parameters, was trained from scratch, and makes a decision in 5 to 7 ms on a consumer GPU, at about $1.81 per million decisions.
- On ZDS-1, a 22,864-case benchmark for decision models, Jev 1.13 scores 78.4, VerQen 73.9 and Laya 52.6.
- VerQen leads by 33 points when the rules are written in the input (93.7% against 60.6%), is 2.0 points behind Jev on real classification with the best calibration of the three, and comes last on open-domain judgement (52.0%).
- Use it for rule-based decisions at volume: policy and eligibility checks, routing, agent guardrails and triage. Do not use it where the answer depends on world knowledge nobody wrote into the input.
- The weights (Apache 2.0), 1.62M training decisions and the ZDS-1 benchmark are public on Hugging Face and GitHub.

- **93.7%** on rules written in the input (Jev 1.13: 60.6%)
- **5.5 ms** per decision, p50 (Jev API: 440 ms)
- **$1.81** per million decisions (Jev API: $28.68)
- **0.039** calibration error on real data (Lowest of the three; lower is better)

## Why a decision model

Most AI products put a chat box in front of a language model. But much of what a product decides all day is not a conversation. Route this ticket. Approve this refund or do not. Is this message a prompt injection. Which tool should the agent call next. Each of those has a fixed set of answers, a rule somewhere that says which one is right, and a cost when the system is confidently wrong.

A chat model can do these jobs, but it does them the expensive way: it writes text you then have to parse, it takes hundreds of milliseconds, and the confidence it reports is not a probability you can set a threshold on. VerQen is built for that job and nothing else. It never writes free text, so there is nothing to parse and nothing to ramble, and its probabilities are calibrated so that "act when above 95%" means what it says.

**State, question, options:** state = Refunds within 30 days for unused items. Delivered 45 days ago, unused.; question = Should the refund be approved?; options = approve | deny | escalate

**Model:** VerQen (232M, 5.5 ms p50)

**Typed decision:** selected = deny_refund; p_correct = 0.67; prediction_set = deny, escalate; should_escalate = true

That example is the real output. VerQen picked the right answer, but its calibrated confidence was below the 95% bar we asked for, so it raised `should_escalate`. That is the intended way to use it: let VerQen act on the decisions it is sure of, and route the rest to a person or a larger model. Three question types are supported: `choice` (one of the options), `noul` (yes or no) and `score` (a point on an ordered scale).

## Run it in a minute

VerQen runs locally, on a GPU or a CPU. Install the package, and the weights download from Hugging Face the first time you load them.

```python
# pip install git+https://github.com/ZorQelis-AI/verqen
from verqen import Verqen

vq = Verqen.from_pretrained("zorqelis-ai/VerQen", risk="0.05")  # act alone only under a 5% error rate
r = vq.decide(
    state="Refund policy: refunds are allowed within 30 days of delivery for unused items.\n"
          "Case: The item was delivered 45 days ago and is unused.",
    question="Should the refund be approved?",
    options=["approve_refund", "deny_refund", "escalate_to_agent"],
)
print(r["selected"], r["p_correct"], r["should_escalate"])  # deny_refund 0.6656 True
```

Every decision returns `selected`, a probability for each option, a calibrated `p_correct`, a conformal `prediction_set` and `should_escalate`. `decide_multi` asks several questions about one state while encoding it once, and `decide_many` batches thousands of decisions a second on one GPU. For high-stakes use, `risk_mode="worst_group"` certifies the error rate on every calibration group rather than on average.

## Results on ZDS-1

To measure a decision model you need a benchmark built for decisions, so we assembled one and published it: the **ZorQelis Decision Suite, ZDS-1**. It has 22,864 cases in three tracks, and it scores every system with one harness, the same inputs and the same pricing rule. We ran VerQen against two other decision models: **Jev 1.13** from TypeSafe, called through OpenRouter's Decisions API, and **Laya** from Convai, an open model we ran on the same GPU as VerQen. The ZDS-1 score is the unweighted mean of the three track accuracies.

*The three tracks of ZDS-1.*

| Track | Cases | What it measures |
| --- | --- | --- |
| **Jev-parity** | 7,840 | Applying a policy written in the input to a case, across 22 business domains: arithmetic, dates, counting, counterfactuals, distractors, prompt injection, missing and conflicting information, long inputs. |
| **Laya-parity** | 14,000 | Real classification and reading tasks from seven public datasets, 2,000 test cases each. |
| **AGY-32** | 1,024 | Open-domain judgement with no rules in the input: 32 domains, 32 cases each, written outside every generator VerQen was trained on. |

*ZDS-1 at a glance. Accuracy %, same harness and inputs for every system. No model wins every track.*

|  | VerQen | Jev 1.13 | Laya |
| --- | --- | --- | --- |
| ZDS-1 score | 73.9 | 78.4 | 52.6 |
| Jev-parity: rules written in the input | 93.7 | 60.6 | 19.3 |
| Laya-parity: real classification sets | 75.9 | 77.9 | 63.1 |
| AGY-32: open-domain judgement, no rules | 52.0 | 96.9 | 75.5 |

*The full leaderboard. Latency is p50 per decision; cost is per million decisions.*

| System | ZDS-1 | Jev-parity | Laya-parity | AGY-32 | Latency | Cost |
| --- | --- | --- | --- | --- | --- | --- |
| Jev 1.13 (API) | **78.4** | 60.6 | **77.9** | **96.9** | 409 to 443 ms | $15 to $29 |
| **VerQen** | 73.9 | **93.7** | 75.9 | 52.0 | **4.7 to 7.3 ms** | **$1.67 to $3.98** |
| Laya | 52.6 | 19.3 | 63.1 | 75.5 | 26 to 58 ms | $6 to $13 |

Jev leads the suite overall, 78.4 to VerQen's 73.9. The overall number hides the shape, though, and the shape is the useful part: the three models are good at different things.

### Where VerQen leads: rules written in the input

On Jev-parity VerQen scores 93.7%, 33 points ahead of Jev. The largest gaps are the cases a production system most needs to get right. When a fact the rule depends on is missing, VerQen asks for it 99.6% of the time; Jev decides anyway and asks 0.7% of the time. When two sources disagree, VerQen flags it 100.0% of the time against Jev's 7.3%. And it follows the policy rather than an instruction hidden in the case 93.0% of the time, against 53.2%.

*Jev-parity, by capability. Accuracy % on the capabilities where the models differ most. All ten are in the table below.*

|  | VerQen | Jev 1.13 | Laya |
| --- | --- | --- | --- |
| Missing information | 99.6 | 0.7 | 1.9 |
| Conflicting sources | 100.0 | 7.3 | 0.6 |
| Prompt injection | 93.0 | 53.2 | 16.2 |
| Arithmetic | 88.2 | 64.3 | 21.8 |
| Long inputs | 86.8 | 83.1 | 26.5 |

*Jev-parity by capability, accuracy %. Best in bold.*

| Capability | VerQen | Jev 1.13 | Laya |
| --- | --- | --- | --- |
| Missing information | **99.6** | 0.7 | 1.9 |
| Conflicting sources | **100.0** | 7.3 | 0.6 |
| Prompt injection | **93.0** | 53.2 | 16.2 |
| Grounded rules | **91.8** | 82.6 | 29.0 |
| Arithmetic | **88.2** | 64.3 | 21.8 |
| Dates | **89.3** | 75.8 | 23.0 |
| Counting | **94.8** | 76.2 | 24.8 |
| Counterfactuals | **91.9** | 87.7 | 34.3 |
| Distractors | **96.1** | 83.5 | 18.1 |
| Long inputs | **86.8** | 83.1 | 26.5 |

> Jev-parity is VerQen's home ground, and we say so plainly. ZorQelis AI built it with the same decision engine that generated part of VerQen's training data. It uses held-out phrasings, but the task format is VerQen's own, and two of VerQen's fixes (for missing and conflicting information) were chosen after seeing its earlier scores on this track.

### Close on real classification, and the best calibrated

On Laya-parity, seven public classification and reading datasets, VerQen scores 75.9%, 2.0 points behind Jev and 13 ahead of Laya. It leads on Banking77, 77-way intent routing, at 85.5%, and on emotion at 91.5%. Jev leads on the tasks that lean on reading comprehension and world knowledge, BoolQ and XNLI.

*Laya-parity by dataset, accuracy %, 2,000 test cases each. Best in bold.*

| Dataset | VerQen | Jev 1.13 | Laya |
| --- | --- | --- | --- |
| Banking77 (77 intents) | **85.5** | 79.5 | 35.8 |
| emotion (6 emotions) | **91.5** | 59.5 | 58.2 |
| AG News (topics) | 89.8 | 89.6 | **93.2** |
| MASSIVE (intents, en) | 77.6 | **80.1** | 43.5 |
| BoolQ (yes or no, from a passage) | 71.2 | **91.8** | 78.0 |
| XNLI (inference, en) | 68.5 | **87.5** | 86.2 |
| SST-5 (5-level sentiment) | 47.4 | **57.2** | 47.0 |

Accuracy is half of what a decision model owes you. The other half is whether its confidence can be trusted, measured here as expected calibration error (lower is better). On Laya-parity VerQen's is 0.039, the lowest of the three. On average its stated confidence sits within 4 points of how often it is actually right, which is what lets the escalation threshold in the quickstart mean what it says.

*Expected calibration error by track. Lower is better; best in bold.*

| Track | VerQen | Jev 1.13 | Laya |
| --- | --- | --- | --- |
| Jev-parity | **0.095** | 0.133 | 0.521 |
| Laya-parity | **0.039** | 0.091 | 0.124 |
| AGY-32 | 0.143 | **0.013** | 0.052 |

### Where VerQen trails: judgement from world knowledge

AGY-32 asks for judgement with no rules in the input, across 32 domains from chess tactics to medical triage. VerQen scores 52.0%, last of the three; Jev scores 96.9% and Laya 75.5%. Its best domains are sentiment analysis (87.5%), content moderation toxicity (84.4%), hr candidate screening (84.4%); its weakest, at 28.1% each, are education student mastery, email urgency scoring, geography navigation routes, math arithmetic reasoning, medical triage acuity.

That is the honest cost of a small model trained from scratch. VerQen learned its language from 7.05B tokens of pretraining; the models it is compared with have seen far more of the world. We included this track because it is the least favourable one for VerQen, and closing it is the next thing we are working on.

## Speed and cost

VerQen reads the whole input once and scores every option in the same pass, so a decision costs one forward pass of a 232M encoder. On a consumer RTX 5060 Ti that is 5.5 ms at the median, about 80 times faster than the Jev API and 16 times cheaper per decision. On an RTX 5090 it is 2.8 ms.

*Latency per decision. p50 in milliseconds on Laya-parity. Jev's latency includes the network round trip to OpenRouter.*

|  | VerQen | Jev 1.13 | Laya |
| --- | --- | --- | --- |
| Median latency | 5.5 ms | 440.1 ms | 26.4 ms |

*Cost per million decisions. US dollars on Laya-parity. VerQen and Laya priced at a cloud L4-class GPU ($0.80 an hour) at their measured latency; Jev as billed.*

|  | VerQen | Jev 1.13 | Laya |
| --- | --- | --- | --- |
| Cost | 1.81 | 28.68 | 5.97 |

For scale: all 22,864 ZDS-1 cases cost $0.64 on the Jev API. That is cheap for a benchmark and expensive for a system that makes a decision on every request, which is the workload VerQen was built for.

## Where to use VerQen

Use VerQen where the rule is known and the decision is frequent: the policy or the playbook is in the input, the answers are a fixed set, and a wrong answer delivered confidently costs more than a slow one. Look elsewhere when the answer depends on broad knowledge nobody wrote down for the model.

| Use VerQen for | Look elsewhere for |
| --- | --- |
| Policy and eligibility checks: refunds, returns, KYC, claims, access | Judgement that rests on broad world knowledge, with no rules given |
| Routing tickets, intents and queues (Banking77: 85.5%) | Languages other than English |
| Agent guardrails: which tool, is it in scope, is it an injection, when to stop | Generating text of any kind |
| Triage at volume, escalating the uncertain cases instead of guessing | High-stakes medical, legal or financial calls without a person reviewing them |

A pattern that works well: put VerQen in front of a larger model. VerQen answers the confident majority in milliseconds, for a fraction of a cent per thousand, and sends only the cases it flags to the slower, more expensive system. You pay large-model prices only where you need large-model knowledge.

## How we built it

VerQen was trained from scratch on two NVIDIA RTX 5090 GPUs in about 23 hours of compute. Every stage is in the public repository, with the data builders, so the training decisions can be regenerated from their seeds.

1. **Pretraining.** Masked-language modelling on a curated 4.07B-token corpus (FineWeb-Edu, Cosmopedia v2, FineMath, Wikipedia and synthetic text) for 7.05B tokens, then context extension to 2,048 tokens. Masked-language modelling beat replaced-token detection in an A/B we registered before running it.
2. **Decision training.** 1.51M typed decisions: 105 procedural specifications across 22 business domains, multi-step environments, and 33 public datasets. Labels are computed from each specification's rules, never written by hand.
3. **RLCD.** Reinforcement Learning from Calibrated Decisions: the exact expected reward over the option set, a proper scoring rule that keeps probabilities honest, and group-relative policy gradients on multi-step environments.
4. **Top-up.** Oracle trajectories with deliberate wrong turns to recover from, 32 phrasings of missing and conflicting facts, cases buried in long unrelated text, and rehearsal of the real datasets. Blended with the earlier weights (WiSE-FT) so no domain regressed.
5. **Format diversity.** Real tasks rewritten in many layouts, field names and label words, so VerQen learns the task rather than the template. In layouts it never saw, XNLI rose from 40.7% to 68.6% and Banking77 from 65.9% to 85.5%.
6. **Calibration.** Temperature scaling, conformal prediction sets and per-group risk control, fitted on a held-out calibration split.

**Every stage had to earn its release.** A candidate replaced the previous model only if it passed a rule fixed before its results were seen: no more than 0.5 points lost overall and no more than 3 points in any domain or question type. Twice a candidate with large gains failed that rule, and it did not ship until the regression behind the failure was fixed.

The bugs worth writing down were the silent ones. The first version learned to recognise the three sentences we used for "this fact is missing", not the idea: on a fourth, held-out wording it flagged a conflict 2.4% of the time. Thirty-two phrasings fixed that. Another was a path that differed between machines, so a data builder quietly found no distractor text and built 1.5 million decisions without the unrelated sentences they were designed to contain. Both are written up in the repository, with what the next run changes because of them.

## ZDS-1, for anyone building a decision model

There was no shared benchmark for decision models, so each one was measured on its own terms. ZDS-1 is our attempt at a common one. It ships with the cases, a harness that scores local models and APIs the same way, the result file behind every number in this post, and a paired comparison tool that reports bootstrap confidence intervals and a McNemar test per capability.

```bash
git clone https://huggingface.co/datasets/zorqelis-ai/zds-1 && cd zds-1
python harness/zds1.py --adapter verqen --checkpoint zorqelis-ai/VerQen --out-dir results/verqen
OPENROUTER_API_KEY=... python harness/zds1.py --adapter jev_openrouter --model typesafe/jev-1.13 --out-dir results/jev
```

Four of the seven Laya-parity datasets have licences that do not let us redistribute them, so the first run downloads them from their original sources and rebuilds the exact cases, and it refuses to score if the rebuilt file does not match the published checksum. VerQen saw the training splits of six of the seven datasets, never a test split, and training rows that exactly matched a test case were removed. Laya's and Jev's training data are not published.

## Open weights, open data, open benchmark

- [zorqelis-ai/VerQen](https://huggingface.co/zorqelis-ai/VerQen): Weights, tokenizer, calibration and the model card. Apache 2.0.
- [verqen-decisions](https://huggingface.co/datasets/zorqelis-ai/verqen-decisions): 1.62M typed decisions in five configurations, with a licence on every row.
- [ZDS-1](https://huggingface.co/datasets/zorqelis-ai/zds-1): 22,864 cases in three tracks, the harness and every result file.
- [ZorQelis-AI/verqen](https://github.com/ZorQelis-AI/verqen): The model, the training pipeline, the data builders and the runbook.

Six datasets VerQen trained on have licences that do not allow redistribution (RACE, SciQ, emotion, tweet_eval sentiment, AG News and SST-5), so they are not in the published training set. The dataset card lists them, with the number of rows used from each, and the build script converts them from their original sources.

## What comes next

VerQen is our first decision model, and ZDS-1 shows exactly where the next one has to improve: knowledge of the world beyond the input, which is AGY-32, and languages beyond English. We will measure the next version on the same suite, publish the result whichever way it goes, and keep the benchmark open for anyone who wants to beat it.

If you build something with VerQen, or find a case it gets confidently wrong, open an issue on [GitHub](https://github.com/ZorQelis-AI/verqen). The cases it gets wrong are the ones we most want to see.

## Questions people ask

### What is VerQen?

VerQen is an open decision model from ZorQelis AI. It reads a situation, a question and a fixed list of options, and returns one option with a calibrated probability and a flag for when to escalate. It has 232M parameters, was trained from scratch, and makes a decision in 5 to 7 ms on a consumer GPU.

### Is VerQen open source?

Yes. The weights and code are released under Apache 2.0 at huggingface.co/zorqelis-ai/VerQen and github.com/ZorQelis-AI/verqen. The training data (1.62M decisions) and the ZDS-1 benchmark are also public on Hugging Face.

### How does VerQen compare with Jev and Laya?

On ZDS-1, Jev 1.13 scores 78.4, VerQen 73.9 and Laya 52.6. VerQen leads by 33 points when the rules are written in the input, is close on real classification with the best calibration, and trails on open-domain judgement. It is about 80 times faster and 16 times cheaper per decision than the Jev API.

### What is ZDS-1?

The ZorQelis Decision Suite is a benchmark for decision models: 22,864 cases in three tracks (Jev-parity, Laya-parity and AGY-32), one harness for local models and APIs, and the published result file behind every score. It is at huggingface.co/datasets/zorqelis-ai/zds-1.

### Can VerQen generate text or chat?

No. VerQen only chooses among the options you give it, answers yes or no, or places a case on a scale. That restriction is what makes it fast, cheap and calibrated. Pair it with a language model when you need text.

### What hardware does VerQen need?

Any recent NVIDIA GPU runs it in milliseconds, and it also runs on a CPU at lower throughput. The median decision takes 4.7 to 7.3 ms on an RTX 5060 Ti, depending on input length, and 2.8 ms on an RTX 5090.
