PAIRED MODEL EVALUATION / 2026-09-26

Which function
fits the request?

A public benchmark slice with three candidate representations, official reference answers, and an interactive audit of every case.

KEV-4B / PINNED RELEASEJEV 1.13.0 / FRESH RUN14 COMPLETE RUNS

[ REPORT / 01 ]

Three ways to describe
the same functions.

The same 1,253 public BFCL v4 multiple and live_multiple cases and official gold function names were used in all three views. Within each view, Kev and Jev received identical state, instruction, candidate order, and candidate text. All three views were specified before this Kev comparison began.

Function-selection accuracy

Name + description

KEV
97.0%
JEV
97.8%

Parameters in prose

KEV
84.5%
JEV
98.6%

Full JSON schema

KEV
97.0%
JEV
98.6%
Candidate viewKev correctKev accuracyJev correctJev accuracyKev p50Jev p50
Name + description1215/1,25396.97%1225/1,25397.77%35.34 ms130.69 ms
Parameters in prose1059/1,25384.52%1236/1,25398.64%65.75 ms134.55 ms
Full JSON schema1216/1,25397.05%1235/1,25398.56%70.65 ms137.22 ms

[ REPORT / 02 ]

Function selection only.

The model chooses one function name from the official candidates. This does not score generated arguments, execution, parallel calls, multi-step behavior, memory, or BFCL's overall leaderboard composite. The official dataset and answers are pinned to Gorilla commit f7cf7359b7ac615a0b294831c5ba2bc95ee4a000.

ViewCandidate representation
Name + descriptionOfficial function name and natural-language description.
Full JSON schemaComplete official function object as compact JSON, including parameters.
Parameters in proseNames, types, descriptions, required markers, allowed values, and defaults in prose.

The prose adapter omits some schema fields. The three views therefore vary both representation and, in the description/prose conditions, information content. Their score differences cannot isolate punctuation or formatting alone. The prose view was post-hoc in the old CLM evaluation, but was specified in advance here.

[ REPORT / 03 ]

Separate the two splits.

ViewSplitCasesKevJev
Name + descriptionmultiple20099.00%99.50%
Name + descriptionlive_multiple105396.58%97.44%
Parameters in prosemultiple20098.00%100.00%
Parameters in proselive_multiple105381.96%98.39%
Full JSON schemamultiple20099.50%100.00%
Full JSON schemalive_multiple105396.58%98.29%

[ REPORT / 03B ]

Confidence is another measurement.

These descriptive metrics use the returned probability distribution. ECE compares top confidence with observed accuracy in 10 fixed, equal-width bins. Multiclass Brier scores the full distribution; lower is better for both.

ViewModelECEBrierCoverage at 0.9Error among acceptedConfident wrong
Name + descriptionKEV0.05420.058374.5%0.1%1
Name + descriptionJEV0.00930.035093.9%0.9%10
Parameters in proseKEV0.10780.252333.8%1.4%6
Parameters in proseJEV0.00530.020695.9%0.3%4
Full JSON schemaKEV0.06570.055872.1%0.1%1
Full JSON schemaJEV0.00600.018895.9%0.2%3

Coverage is the share with top probability at least 0.9. Error among accepted divides wrong accepted answers by all accepted answers; confident wrong is the count. The 0.9 threshold is a descriptive display choice, not a production gate fitted or validated on independent data. No model was recalibrated using these evaluation labels. One Jev prose response rounded its probabilities to two decimal places, summing to 0.99; its returned values were preserved without renormalization.

[ REPORT / 04 ]

Inspect every paired answer.

Loading frozen results...

[ REPORT / 05 ]

A fixed historical reference.

The request hashes and source-file hashes match the historical CLM evaluation in every view. The rows below compare the same adapted task. CLM was not rerun. These values also must not be equated with the CLM team's published 95.2% BFCL figure: its exact case set and scoring adapter were not available in the earlier report.

Candidate viewHistorical CLMNew KevNew Jev
Name + description76.86%96.97%97.77%
Parameters in prose58.42%84.52%98.64%
Full JSON schema47.49%97.05%98.56%