Data & evaluation

The data behind
the decisions.

How Vej was trained, what it was tested on, how it compares with Jev, and where it falls short. We kept the receipts.

100,002Training
28,572Validation
14,286Final test

142,860 decisions · 70 / 20 / 10 split · Correlated projections, not unique conversations

Ten sources. Explicit roles.

Counts represent Vej decision rows after task selection, formatting, and screening. External links connect directly to the original academic and commercial publishers. Source texts retain their original licences.

On narrow screens, scroll the source table horizontally.

Source decisions by Vej split role and original licence
Source Training Validation Test Terms
WANLI 6,144 1,024 0 CC BY 4.0
CLINC150/OOS 7,424 1,024 0 CC BY 3.0
Vej workflows 3,072 384 0 Apache-2.0
HelpSteer2 9,000 800 0 CC BY 4.0
MultiNLI 30,000 9,000 0 Selected OANC terms
SNLI 2,048 512 0 CC BY-SA 4.0
HH-RLHF 15,000 2,500 0 MIT
BANKING77 3,072 384 0 CC BY 4.0
HelpSteer1 8,000 4,000 5,000 CC BY 4.0
Schema-Guided Dialogue 16,242 8,944 9,286 CC BY-SA 4.0

No raw dataset payloads are hosted on this website. Vej's Apache-2.0 licence does not relicense source texts. MultiNLI fiction genres are excluded.

Checkpoint selection

How the checkpoint was trained.

Vej trains rank-16, alpha-32 LoRA adapters and a native FP32 scalar decision head on a frozen Qwen3.5-2B text backbone (revision 15852e8). The original generative pretraining objective does not determine Vej's decision-oriented output interface.

  1. Public pilot. 4,096 training and 256 development decisions. Two epochs, 512 updates, selected at step 256.
  2. Expanded training. 32,768 training and 2,048 development decisions. Two epochs, selected at step 1,280 with validation NLL 0.636590.
  3. First 100k stage. 100,002 training and 28,572 validation decisions. An unexpected system reboot interrupted training. The saved selected checkpoint was step 1,024 with validation NLL 0.492022.
  4. Recovery stage. Initialized with a fresh optimizer state, seed 43, across one shuffled epoch of the same corpus (6,251 updates). Step 6,144 was selected by lowest pooled validation NLL of 0.458497.
  5. Final evaluation. The selected checkpoint was tested against all 14,286 frozen test decisions on 26 September 2026.

Recovery repeated example exposure rather than performing an exact optimizer resume. Training stages overlap, so their counts cannot be summed to claim unique volume.

Vej workflows represent authored cases covering customer support, security, billing, and agent traces. Public sources supply natural language inference relations, intent classifications, human preference comparisons, and task dialogue state.

Roles & exposure

A split is a protocol.

A split is a protocol for how each slice of data may be used, not just three files. Training updates parameters, validation selects checkpoints, and the final test was evaluated once.

The test set holds 5,000 HelpSteer1 ratings and 9,286 Schema-Guided Dialogue decisions. Both test sources also contribute training data, so this measures domain transfer rather than an unseen benchmark.

Known exposure includes twelve inherited WANLI validation decisions, 19 historical training decisions, and one inspected comparator example.

Final test / Qwen 2B alpha2

What was measured.

Exact label accuracy on 14,286 frozen decisions evaluated on 26 September 2026. This test measures categorical decision accuracy, not end-to-end tool execution.

Vej alpha2 final test exact accuracy by primitive
Question type Test Decisions Exact Accuracy
Choice 6,197 89.22%
Noul 3,089 91.71%
Score (exact rubric level) 5,000 51.62%
Overall 14,286 76.60%

Score needs more work.

51.62% exact-level accuracy against 48.44% for a post-hoc test-majority control that uses the test labels.

Tool readiness is not permission.

Development suites passed 30 of 48 harder tool cases, with nine wrong calls among 31 "ready" predictions. Details below.

Calibration and confidence distribution

Across the final test, overall negative log-likelihood is 0.589448, summed Brier score is 0.313870, and descriptive ten-bin expected calibration error (ECE) is 0.053473. No post-hoc calibration was applied. Decisions with maximum probability ≥ 0.9 cover 58.48% of the test set with an error rate of 5.53%. The 0.8 to 0.9 confidence bin shows 59.22% accuracy at 85.16% mean confidence.

Full aggregate evaluation JSON   Descriptive controls

Hosted comparison / Same test rows

Vej and Jev on the same rows.

Jev is TypeSafe's hosted decision model and API. To evaluate how Vej compares with Jev's decision contract, hosted Jev 1.13 was tested against the exact same 14,286 held-out decisions. Across the full sample, Vej achieved 76.60% exact accuracy (10,943 correct) compared to 71.76% (10,251 correct) for hosted Jev 1.13 using Jev's documented post-launch rounding adjustment. Vej leads overall and on Choice and Score. Jev leads on Noul. The 4.84 percentage point difference reflects this fixed test sample, not a claim of general superiority, broader transfer, or statistical significance.

Same-row exact accuracy on 14,286 decisions. Jev uses the rounding-adjusted supplement.

Scroll horizontally to see the full comparison.

Vej and Jev same-row exact accuracy
Primitive Rows Vej alpha2 Hosted Jev 1.13 Difference (points)
Choice 6,197 89.22% 83.07% Vej (+6.15)
Noul 3,089 91.71% 93.65% Jev (+1.94)
Score (exact level) 5,000 51.62% 44.20% Vej (+7.42)
Overall 14,286 76.60% 71.76% Vej (+4.84)

How it was compared

The comparison matched dataset hash, row order, question identifiers, and candidate ordering. Hosted Jev returned 14,188 strictly valid responses (71.95% vs 76.80% for Vej on those rows), while 98 non-normalised maps summed to 0.99 due to two-decimal output precision. This is one fixed test, not Jev's own benchmark. Speed and cost were not compared.

Aggregate comparison and evidence hashes

Development diagnostics

Where tool use falls short.

  • Across 64 complete workflow actions, Vej completed 60 actions correctly, but made two critical errors where the wrong action was selected.
  • On 48 harder tool diagnostic cases, Vej completed 30 cases, with nine incorrect calls among 31 ready predictions.
  • On 12 original tool cases, Vej completed 4 cases, with three incorrect calls among seven ready predictions.

These are development suites, not a held-out test, and they do not measure real tool execution.

Aggregate tool diagnostic evidence

Frozen identities

Hashes you can check. Bring a terminal.

Evaluation was conducted on an AMD Radeon RX 9070 XT 16 GiB workstation with PyTorch 2.14 and ROCm 7.2.

Weights · SHA256
b99188eedcbc6e1faab966ed2c827025b29360157298d0cc86e53a103d73bf96

Training · SHA256
4b338fa71c6ec02161ad698c3f3789672897cbed4dd036f6a822f6cc6c38fea6

Validation · SHA256
488de164255ee1e22334a844051b1e837d5f8afa5cc7792159c93a7838ad53bc

Final test · SHA256
ac1b11d62e55c6094101f71790f5bcd326b1f32afb91311bb334238036311265

No individual example predictions or raw training texts are hosted on this website.