Model card / Research alpha

Run Vej
yourself.

Vej Qwen 2B alpha2 returns native probabilities over outcomes you supply. It runs on your own GPU, and your code applies the policy.

Architecture

A decision head, not text generation.

Qwen3.5-2B provides the pretrained text backbone. Vej freezes its base BF16 parameters and trains rank-16, alpha-32 LoRA adapters alongside a custom FP32 scalar decision head. There is no vocabulary projection head or token generator in this text-only checkpoint. This checkpoint makes text decisions; the codebase also carries an untrained vision input path, and no vision decision quality has been measured.

State + question + outcomeScalar score

State, question, and candidate text are concatenated and passed through the model to produce a scalar score. Candidate scores are normalised into probabilities within each independent question, with zero-based rubric expectations computed in application code.

Base model
Qwen/Qwen3.5-2B, revision 15852e8
Text backbone
Approximately 1.88 billion parameters (frozen BF16)
Trainable parameters
Approximately 16.82 million (LoRA adapters + scalar head)
Checkpoint size
Full native state dict, approximately 3.6 GiB
Input token limit
2,048 tokens per rendered candidate by default (overlong inputs fail closed). The native Qwen loader accepts --max-length 49152; long-context quality is not validated.
Selected checkpoint
Recovery step 6,144 (lowest pooled validation NLL: 0.458497)
Execution mode
Independent candidate batching with repeated state encoding

Availability

Privately published on Hugging Face.

Access to the checkpoint is by request. Contact lode@lodekennes.com.

The checkpoint is hosted at Kennes/vej-qwen-2b-alpha2 on Hugging Face as a private model for authorised account members. The package contains the native checkpoint, tokenizer assets, source code snapshot, metrics, and checksums.

The core development repository at LodeKennes/Vej is also private. This public website is hosted from LodeKennes/vej-site.

The experimental classification export is also private on Hugging Face: Kennes/vej-qwen-2b-alpha2-vllm.

Local inference

Run it on your GPU.

Authorised users can download the release package with their authorised Hugging Face account, verify SHA256SUMS, extract vej-source.tar.gz, and run predictions directly from that source directory. Python 3.12 to 3.14 is supported.

uv sync --frozen --extra cpu --extra qwen
HF_HUB_OFFLINE=1 uv run --frozen --extra cpu --extra qwen python -m vej predict \
  examples/decisions.json --checkpoint /path/to/download/checkpoint --device cpu

CPU execution is a functional fallback and can be slow. The backbone remains BF16. Disabling autocast does not convert it to FP32. For compatible CUDA or ROCm environments, install the source package with its qwen extra in your GPU environment, then execute:

HF_HUB_OFFLINE=1 python -m vej predict examples/decisions.json \
  --checkpoint /path/to/download/checkpoint --device cuda \
  --precision bfloat16 --candidate-batch-size 4

Local HTTP and SDKs

vej serve --backend local --checkpoint /path/to/download/checkpoint --device cuda

The server binds to 127.0.0.1 by default, without authentication. Keep it local. It exposes POST /v1/predict, /healthz, /readyz and /metrics; native responses include usage.processed_tokens.

The official TypeSafe (Jev) Python SDK uses POST /v1/systemone. Cloudflare Clef uses POST /client/v4/accounts/{account_id}/ai/run/@cf/cloudflare/{clef|clef-flash}. Both text wire formats work on local and experimental vLLM backends.

The serving commands, SDK adapters, vLLM exporter and context override require the current private Vej source with its qwen runtime dependencies. The source archive bundled with the alpha2 release predates these features; the prediction commands above work with that archive.

Experimental vLLM export

vej export-vllm writes a self-contained classification export for vLLM ≥ 0.27, using --runner pooling --convert classify. The private published export can also be served without those flags. The recorded alpha2 runs have not passed the numerical and contract parity gates; this remains experimental, not a supported serving path. Use Vej’s native loader for supported inference.

Intended use

Research decision support.
Known limits.

The final test recorded 76.60% exact accuracy across 14,286 decisions. Score accuracy and tool diagnostics remain weak, confidence is not calibrated, and probabilities are not permission to act. Keep independent policy checks around consequential actions; the same-test comparison with hosted Jev is on the evaluation page.

Vej code and model adaptations use Apache-2.0, while Qwen base assets and training datasets retain their respective terms. See licences and attribution.