Vej / Native decision model
State in,
probabilities
out.
Vej is a 2B text model you run on your own GPU. Describe a situation and the outcomes that matter; Vej returns a probability for each. No prose to parse. Call it through the CLI, local HTTP, or Jev/Clef SDK wire formats. Your code makes the call.
Qwen 2B alpha2 · Text only · Private checkpoint
A native decision
What should happen next?
Illustrative values. No model runs in your browser.
The problem
Chat models answer.
Your workflow needs a decision.
01 / PARSING
Prose you have to parse.
A yes, a category or a rating comes back wrapped in sentences or loosely shaped JSON. So you add schema prompts, validators and retries, and the workflow still stops when one reply arrives wearing markdown.
02 / CONFIDENCE
A confidence that is not a probability.
Ask how sure it is and it writes a number. Nothing makes those numbers sum to one or hold steady between calls. You cannot trust a threshold.
03 / CONTROL
Data that leaves your process.
Each decision sends internal state to an endpoint you do not control. Run Vej inside your own infrastructure, or, in future, on a planned EU service designed not to retain your content.
How it works
Three steps from
situation to probabilities.
One request carries up to 128 independent questions, each scored on its own.
01 / DESCRIBE
Send the state.
Pass the situation as text or JSON: a message, record, or document.
02 / ASK
Ask independent questions.
Choice picks among named options. Noul tests whether a statement holds. Score evaluates rubric levels. Questions do not see each other.
03 / ACT
Apply your policy.
Each answer sums to one. Your code sets thresholds and gates actions. Vej never executes anything.
REQUEST · examples/decisions.json
{
"state": {
"message": "Please cancel my subscription. I have not provided an account ID.",
"account_id": null,
"account_verified": false
},
"questions": {
"route": {
"type": "choice",
"question": "Which action best matches the request?",
"options": {
"cancel_subscription": "Cancel an existing subscription",
"lookup_invoice": "Retrieve a billing invoice",
"none": "Neither offered action applies"
}
},
"can_execute": {
"type": "noul",
"statement": "The account is verified and its ID is available."
},
"completeness": {
"type": "score",
"question": "How complete is the information needed to execute cancellation?",
"levels": [
"The account ID and account verification are both missing",
"Exactly one of account ID or account verification is missing",
"The account ID is available and the account is verified"
]
}
}
}
ILLUSTRATIVE RESPONSE · abridged
{
"answers": {
"route": {
"type": "choice",
"choice": "cancel_subscription",
"probabilities": {
"cancel_subscription": 0.91,
"lookup_invoice": 0.05,
"none": 0.04
}
},
"can_execute": {
"type": "noul",
"noul": 0.02,
"probabilities": { "true": 0.02, "false": 0.98 }
},
"completeness": {
"type": "score",
"score": 0.14,
"probabilities": [0.88, 0.10, 0.02]
}
}
}
Why Vej
Numbers your code can use.
Hardware you already own.
Decisions as distributions.
Every question returns probabilities summing to one over your outcomes. Route, gate, and score without parsing a word.
Runs where your data lives.
Trained and evaluated on a single 16 GB consumer GPU. CPU works when speed does not matter. At inference nothing leaves your machine.
Published aggregate evidence.
Every number here has a published aggregate file and a pinned hash. Weights and splits are private; the weight hash is for authorised users. All ten sources are listed with licence and role.
Apache-2.0 code and adaptations.
Vej code and adaptations are Apache-2.0. Base model and datasets retain their own terms.
Where it fits
Four places a distribution
beats a paragraph.
Route a request.
A support message could mean billing, security, or cancellation. Vej returns probabilities for each route so clear cases dispatch and split cases escalate.
Gate a tool call.
A refund needs a verified account ID. Ask Noul whether the requirement holds and block the call on a confident no.
Score a reply against a rubric.
Grade answers against your rubric. Score returns a distribution over levels, and your code computes the expected position.
Decide when to escalate.
A workflow hits conflicting instructions. Ask Choice to weigh retry, fallback and human handoff, then apply your policy threshold.
Evaluation / 26 September 2026
Tested on 14,286 held-out decisions.
Compared with Jev on the same rows.
HelpSteer1 ratings and Schema-Guided Dialogue decisions, frozen before the final run.
Scroll horizontally to see the full comparison.
| Task | Decisions | Vej alpha2 | Hosted Jev 1.13¹ |
|---|---|---|---|
| Choice · pick an outcome | 6,197 | 89.22% | 83.07% |
| Noul · assess a statement | 3,089 | 91.71% | 93.65% |
| Score · exact rubric level | 5,000 | 51.62% | 44.20% |
| Overall | 14,286 | 76.60% | 71.76% |
¹ Jev is TypeSafe's hosted decision model, shown with its documented rounding adjustment.
Score needs more work.
51.62% exact-level accuracy against 48.44% for a post-hoc majority control that uses the test labels.
Tool readiness is not permission.
A development suite passed 30 of 48 harder tool cases. Probabilities are not permission to act; keep execution behind your own checks.
Status / October 2026
Ways to run. Work still to do.
With the private checkpoint, use the CLI, a local HTTP server, or the Jev and Cloudflare Clef SDK wire formats. An experimental vLLM export also exists; its numerical and contract parity gates have not passed. See local inference.
Research status
The Noul improvement programme paused on 9 October without promoting a new model. No candidate beat alpha2's Noul development reference while also passing the checks that protect its other results. Alpha2 and its published evaluation remain unchanged.
A hosted option is in development.
We are building an EU-hosted API for Choice, Noul and Score: a realtime endpoint, Jev and Clef wire formats, and batch jobs that read from and write to your own EU object storage. It is designed so that request content is not stored on our side; that remains a design goal until the launch checks pass. It will not replace self-hosting. Draft data processing agreement.
Questions
Straight answers.
What Vej returns, how it compares, and what it does not claim.
- Is there a hosted API?
- Not today. An EU-hosted option is in development, designed so that request content is not stored on our side. That is a design goal until the launch checks pass.
- What does Vej return?
- A probability for each outcome you define, within each independent question. Choice picks named outcomes, Noul assesses statements, and Score rates rubric levels.
- How is Vej different from Jev?
- Jev is TypeSafe's hosted decision model and API. Vej offers the same three question types as a 2B model you run yourself, and scored 76.60% against hosted Jev's 71.76% on the same 14,286-decision test. This is not a claim of general superiority.
- Can I run Vej on my own hardware?
- Yes. Inference runs locally through Vej's native loader on a GPU, or on CPU when speed does not matter. The checkpoint is about 3.6 GiB and was trained and evaluated on one 16 GB card.
- Is Vej open source?
- Vej code and adaptations are Apache-2.0, while the base model and datasets keep their own terms. Weights are private on Hugging Face as Kennes/vej-qwen-2b-alpha2 (private). Access to the checkpoint is by request. Contact lode@lodekennes.com.
- Is Vej production-ready?
- No. Vej is a text-only research alpha. Score accuracy and tool diagnostics remain weak, confidence is not calibrated, and no serving-speed claim is made.
- What does “Noul” mean?
- Noul assesses whether a statement holds. It is an intentional product term, not a typo.
- Why the pea pod?
- A pod holds a small number of outcomes you can count, which is the whole idea.
Start with the evidence.
Then the model card. Then, if you hold the checkpoint, your GPU.