Data & evaluation
The data behind
the decisions.
How Vej was trained, what it was tested on, how it compares with Jev, and where it falls short. We kept the receipts.
142,860 decisions · 70 / 20 / 10 split · Correlated projections, not unique conversations
Ten sources. Explicit roles.
Counts represent Vej decision rows after task selection, formatting, and screening. External links connect directly to the original academic and commercial publishers. Source texts retain their original licences.
On narrow screens, scroll the source table horizontally.
| Source | Training | Validation | Test | Terms |
|---|---|---|---|---|
| WANLI | 6,144 | 1,024 | 0 | CC BY 4.0 |
| CLINC150/OOS | 7,424 | 1,024 | 0 | CC BY 3.0 |
| Vej workflows | 3,072 | 384 | 0 | Apache-2.0 |
| HelpSteer2 | 9,000 | 800 | 0 | CC BY 4.0 |
| MultiNLI | 30,000 | 9,000 | 0 | Selected OANC terms |
| SNLI | 2,048 | 512 | 0 | CC BY-SA 4.0 |
| HH-RLHF | 15,000 | 2,500 | 0 | MIT |
| BANKING77 | 3,072 | 384 | 0 | CC BY 4.0 |
| HelpSteer1 | 8,000 | 4,000 | 5,000 | CC BY 4.0 |
| Schema-Guided Dialogue | 16,242 | 8,944 | 9,286 | CC BY-SA 4.0 |
No raw dataset payloads are hosted on this website. Vej's Apache-2.0 licence does not relicense source texts. MultiNLI fiction genres are excluded.
Checkpoint selection
How the checkpoint was trained.
Vej trains rank-16, alpha-32 LoRA adapters and a native FP32 scalar decision head on a frozen Qwen3.5-2B text backbone (revision 15852e8). The original generative pretraining objective does not determine Vej's decision-oriented output interface.
- Public pilot. 4,096 training and 256 development decisions. Two epochs, 512 updates, selected at step 256.
- Expanded training. 32,768 training and 2,048 development decisions. Two epochs, selected at step 1,280 with validation NLL 0.636590.
- First 100k stage. 100,002 training and 28,572 validation decisions. An unexpected system reboot interrupted training. The saved selected checkpoint was step 1,024 with validation NLL 0.492022.
- Recovery stage. Initialized with a fresh optimizer state, seed 43, across one shuffled epoch of the same corpus (6,251 updates). Step 6,144 was selected by lowest pooled validation NLL of 0.458497.
- Final evaluation. The selected checkpoint was tested against all 14,286 frozen test decisions on 26 September 2026.
Recovery repeated example exposure rather than performing an exact optimizer resume. Training stages overlap, so their counts cannot be summed to claim unique volume.
Vej workflows represent authored cases covering customer support, security, billing, and agent traces. Public sources supply natural language inference relations, intent classifications, human preference comparisons, and task dialogue state.
Roles & exposure
A split is a protocol.
A split is a protocol for how each slice of data may be used, not just three files. Training updates parameters, validation selects checkpoints, and the final test was evaluated once.
The test set holds 5,000 HelpSteer1 ratings and 9,286 Schema-Guided Dialogue decisions. Both test sources also contribute training data, so this measures domain transfer rather than an unseen benchmark.
Known exposure includes twelve inherited WANLI validation decisions, 19 historical training decisions, and one inspected comparator example.
Final test / Qwen 2B alpha2
What was measured.
Exact label accuracy on 14,286 frozen decisions evaluated on 26 September 2026. This test measures categorical decision accuracy, not end-to-end tool execution.
| Question type | Test Decisions | Exact Accuracy |
|---|---|---|
| Choice | 6,197 | 89.22% |
| Noul | 3,089 | 91.71% |
| Score (exact rubric level) | 5,000 | 51.62% |
| Overall | 14,286 | 76.60% |
Score needs more work.
51.62% exact-level accuracy against 48.44% for a post-hoc test-majority control that uses the test labels.
Tool readiness is not permission.
Development suites passed 30 of 48 harder tool cases, with nine wrong calls among 31 "ready" predictions. Details below.
Calibration and confidence distribution
Across the final test, overall negative log-likelihood is 0.589448, summed Brier score is 0.313870, and descriptive ten-bin expected calibration error (ECE) is 0.053473. No post-hoc calibration was applied. Decisions with maximum probability ≥ 0.9 cover 58.48% of the test set with an error rate of 5.53%. The 0.8 to 0.9 confidence bin shows 59.22% accuracy at 85.16% mean confidence.
Hosted comparison / Same test rows
Vej and Jev on the same rows.
Jev is TypeSafe's hosted decision model and API. To evaluate how Vej compares with Jev's decision contract, hosted Jev 1.13 was tested against the exact same 14,286 held-out decisions. Across the full sample, Vej achieved 76.60% exact accuracy (10,943 correct) compared to 71.76% (10,251 correct) for hosted Jev 1.13 using Jev's documented post-launch rounding adjustment. Vej leads overall and on Choice and Score. Jev leads on Noul. The 4.84 percentage point difference reflects this fixed test sample, not a claim of general superiority, broader transfer, or statistical significance.
Same-row exact accuracy on 14,286 decisions. Jev uses the rounding-adjusted supplement.
Scroll horizontally to see the full comparison.
| Primitive | Rows | Vej alpha2 | Hosted Jev 1.13 | Difference (points) |
|---|---|---|---|---|
| Choice | 6,197 | 89.22% | 83.07% | Vej (+6.15) |
| Noul | 3,089 | 91.71% | 93.65% | Jev (+1.94) |
| Score (exact level) | 5,000 | 51.62% | 44.20% | Vej (+7.42) |
| Overall | 14,286 | 76.60% | 71.76% | Vej (+4.84) |
How it was compared
The comparison matched dataset hash, row order, question identifiers, and candidate ordering. Hosted Jev returned 14,188 strictly valid responses (71.95% vs 76.80% for Vej on those rows), while 98 non-normalised maps summed to 0.99 due to two-decimal output precision. This is one fixed test, not Jev's own benchmark. Speed and cost were not compared.
Development diagnostics
Where tool use falls short.
- Across 64 complete workflow actions, Vej completed 60 actions correctly, but made two critical errors where the wrong action was selected.
- On 48 harder tool diagnostic cases, Vej completed 30 cases, with nine incorrect calls among 31 ready predictions.
- On 12 original tool cases, Vej completed 4 cases, with three incorrect calls among seven ready predictions.
These are development suites, not a held-out test, and they do not measure real tool execution.
Frozen identities
Hashes you can check. Bring a terminal.
Evaluation was conducted on an AMD Radeon RX 9070 XT 16 GiB workstation with PyTorch 2.14 and ROCm 7.2.
Weights · SHA256b99188eedcbc6e1faab966ed2c827025b29360157298d0cc86e53a103d73bf96
Training · SHA2564b338fa71c6ec02161ad698c3f3789672897cbed4dd036f6a822f6cc6c38fea6
Validation · SHA256488de164255ee1e22334a844051b1e837d5f8afa5cc7792159c93a7838ad53bc
Final test · SHA256ac1b11d62e55c6094101f71790f5bcd326b1f32afb91311bb334238036311265
No individual example predictions or raw training texts are hosted on this website.