ZorQelis AI: Models that work a problem through.

Most assistants answer from the surface of a question. We build for the harder case: holding a problem across many steps, deciding what to do next, and noticing when an earlier step was wrong. The SoreQen series is where that work ships.

The SoreQen S1 series

Three sizes, one training recipe. Deliberately small: cheap enough to offer generously and fast enough to feel immediate, rather than large enough to top a leaderboard.

SoreQen S1 Mini
0.8B parameters

Fastest. Good for quick questions and everyday chat.

SoreQen S1
2B parameters

The default. Balanced speed and depth.

SoreQen S1 Mega
4B parameters

Deepest reasoning. Best for hard problems and long documents.

All three take the full capability set: vision, tool calling, structured output, reasoning at every setting, and a 256k context window. There is no tier that unlocks a feature. The models in detail →

VerQen, the decision model

Beyond the chat box. VerQen reads a situation, a question and a fixed list of options, and returns one calibrated decision in about 5 milliseconds, with a flag for when to hand over. Open weights, open data, and a published benchmark that includes the tracks it loses.

Built for rule-based decisions at volume: policy checks, routing, agent guardrails and triage. VerQen in detail → The launch post →

Open by default

The weights are published, and so is the data they were trained on. Both, because weights alone let you run a model without letting you check it.

4
open-weight models: three SoreQen S1 and VerQen
600,229
SoreQen training rows, published
3
SoreQen training datasets, with their filtering statistics

As far as we know, the Hinglish set is the only curated public instruction corpus of its kind. Weights, data and licences →

What it cost to get there

The S1 models were trained in a single month. Forty-five things went wrong, and the ones worth writing down were the ones that went wrong silently.

A run that crashes teaches you something immediately. A run that completes while doing the wrong thing teaches you nothing until you check, and the checkpoint that was quietly training from random initialisation, in a run whose loss curve looked entirely ordinary, is the reason we now diff every checkpoint against its base.

The postmortem is published because a company that only reports its results is asking to be taken on faith. What went wrong, and why →

Where to use it

From the blog