Which function
fits the request?
A public benchmark slice with three candidate representations, official reference answers, and an interactive audit of every case.
[ REPORT / 01 ]
Three ways to describe
the same functions.
The same 1,253 public BFCL v4 multiple and live_multiple cases and official gold function names were used in all three views. Within each view, Kev and Jev received identical state, instruction, candidate order, and candidate text. All three views were specified before this Kev comparison began.
Function-selection accuracy
Name + description
Parameters in prose
Full JSON schema
| Candidate view | Kev correct | Kev accuracy | Jev correct | Jev accuracy | Kev p50 | Jev p50 |
|---|---|---|---|---|---|---|
| Name + description | 1215/1,253 | 96.97% | 1225/1,253 | 97.77% | 35.34 ms | 130.69 ms |
| Parameters in prose | 1059/1,253 | 84.52% | 1236/1,253 | 98.64% | 65.75 ms | 134.55 ms |
| Full JSON schema | 1216/1,253 | 97.05% | 1235/1,253 | 98.56% | 70.65 ms | 137.22 ms |
[ REPORT / 02 ]
Function selection only.
The model chooses one function name from the official candidates. This does not score generated arguments, execution, parallel calls, multi-step behavior, memory, or BFCL's overall leaderboard composite. The official dataset and answers are pinned to Gorilla commit f7cf7359b7ac615a0b294831c5ba2bc95ee4a000.
| View | Candidate representation |
|---|---|
| Name + description | Official function name and natural-language description. |
| Full JSON schema | Complete official function object as compact JSON, including parameters. |
| Parameters in prose | Names, types, descriptions, required markers, allowed values, and defaults in prose. |
The prose adapter omits some schema fields. The three views therefore vary both representation and, in the description/prose conditions, information content. Their score differences cannot isolate punctuation or formatting alone. The prose view was post-hoc in the old CLM evaluation, but was specified in advance here.
[ REPORT / 03 ]
Separate the two splits.
| View | Split | Cases | Kev | Jev |
|---|---|---|---|---|
| Name + description | multiple | 200 | 99.00% | 99.50% |
| Name + description | live_multiple | 1053 | 96.58% | 97.44% |
| Parameters in prose | multiple | 200 | 98.00% | 100.00% |
| Parameters in prose | live_multiple | 1053 | 81.96% | 98.39% |
| Full JSON schema | multiple | 200 | 99.50% | 100.00% |
| Full JSON schema | live_multiple | 1053 | 96.58% | 98.29% |
[ REPORT / 03B ]
Confidence is another measurement.
These descriptive metrics use the returned probability distribution. ECE compares top confidence with observed accuracy in 10 fixed, equal-width bins. Multiclass Brier scores the full distribution; lower is better for both.
| View | Model | ECE | Brier | Coverage at 0.9 | Error among accepted | Confident wrong |
|---|---|---|---|---|---|---|
| Name + description | KEV | 0.0542 | 0.0583 | 74.5% | 0.1% | 1 |
| Name + description | JEV | 0.0093 | 0.0350 | 93.9% | 0.9% | 10 |
| Parameters in prose | KEV | 0.1078 | 0.2523 | 33.8% | 1.4% | 6 |
| Parameters in prose | JEV | 0.0053 | 0.0206 | 95.9% | 0.3% | 4 |
| Full JSON schema | KEV | 0.0657 | 0.0558 | 72.1% | 0.1% | 1 |
| Full JSON schema | JEV | 0.0060 | 0.0188 | 95.9% | 0.2% | 3 |
Coverage is the share with top probability at least 0.9. Error among accepted divides wrong accepted answers by all accepted answers; confident wrong is the count. The 0.9 threshold is a descriptive display choice, not a production gate fitted or validated on independent data. No model was recalibrated using these evaluation labels. One Jev prose response rounded its probabilities to two decimal places, summing to 0.99; its returned values were preserved without renormalization.
[ REPORT / 04 ]
Inspect every paired answer.
Loading frozen results...
[ REPORT / 05 ]
A fixed historical reference.
The request hashes and source-file hashes match the historical CLM evaluation in every view. The rows below compare the same adapted task. CLM was not rerun. These values also must not be equated with the CLM team's published 95.2% BFCL figure: its exact case set and scoring adapter were not available in the earlier report.
| Candidate view | Historical CLM | New Kev | New Jev |
|---|---|---|---|
| Name + description | 76.86% | 96.97% | 97.77% |
| Parameters in prose | 58.42% | 84.52% | 98.64% |
| Full JSON schema | 47.49% | 97.05% | 98.56% |