Kev / Jev.
The same questions.
A direct follow-up to the CLM comparison: two decision models, the original benchmarks, fresh results, and visible limits.
[ REPORT / 00 ]
The same tests.
Two new sets of answers.
A new paired comparison of the public Kev-4B checkpoint and hosted TypeSafe Jev uses the original CLM evaluation fixtures and game rules. Both models received identical fixed requests; each played the same five game seeds independently.
The model
drives.
Five 60-second courses per model, with and without the deterministic safety shield.
The model
advises.
369 narrow decisions and 40 support tickets, including an exact replay.
The model
routes a tool.
1,253 public BFCL cases in three candidate representations. Inspect every paired answer.
The results
in words.
Numbered facts, notable observations, setup details, and a conclusion.
[ REPORT / 05 ]
Setup and audit trail.
Kev-4B is pinned to 139fdd94f1b6, served over Qwen/Qwen3.5-4B-Base in bfloat16, with its shipped temperature of 2.41, fused GPU kernels, CUDA graphs, and the default prefix cache. No fine-tuning, workload calibration, or date preprocessing was applied.
Both clients ran on one US-KS-2 GPU host with an NVIDIA RTX PRO 4500 Blackwell Server Edition (32 GB). Kev used loopback HTTP; Jev used hosted HTTPS and identified itself as jev-1.13.0. Connections and common question shapes were warmed before each suite. A fresh Kev server was started before each BFCL view. No client answer cache was used.
Response times include these different network paths, serving caches, and compute stacks. They measure the deployed setups, not isolated model compute. The full raw results and environment details are available below.