Tool-call parser conformance matrix
Each cell shows how one inference engine's own tool-call and reasoning parsers handle one model family's recorded outputs, parsed once without streaming and once per token-chunking strategy. Pass rates are strict; whitespace-only differences are counted separately as soft passes. Latest run: .
- pass strict match on every realistic strategy
- soft pass only whitespace differs (normalization
soft-v1) - fail a check failed
- error the harness failed, not the engine's parser
- unsupported the engine has no parser for this family or model
Families × engines
| Family | llamacpp a25c9865 run | ollama 7af39318 run | sglang 0.5.20 run | transformers 5.17.0 run | vllm 0.30.0 run |
|---|---|---|---|---|---|
DeepSeek (V3/R1, V3.1, V3.2 DSML, V4 DSML, V4.1 DSML)
deepseek |
fail
55%
35 pass · 29 fail · 21 unsupported
|
pass 100% 12 pass · 73 unsupported |
fail
20%
13 pass · 23 soft pass · 28 fail · 21 unsupported
|
unsupported 85 unsupported |
fail
51%
43 pass · 9 soft pass · 33 fail
|
Gemma 4 (<|tool_call>call:NAME{...} object notation)
gemma4 |
fail
73%
35 pass · 5 soft pass · 8 fail
|
fail
83%
40 pass · 1 soft pass · 7 fail
|
fail
62%
30 pass · 5 soft pass · 13 fail
|
fail
69%
33 pass · 7 soft pass · 8 fail
|
fail
62%
30 pass · 2 soft pass · 16 fail
|
GLM (4.5, 4.6, 4.7, 5.x)
glm |
fail
80%
41 pass · 2 soft pass · 8 fail
|
fail
91%
29 pass · 3 fail · 19 unsupported
|
fail
49%
25 pass · 2 soft pass · 24 fail
|
unsupported 51 unsupported |
fail
65%
33 pass · 11 soft pass · 7 fail
|
gpt-oss (Harmony response format)
gpt-oss |
fail
84%
43 pass · 8 fail
|
fail
88%
45 pass · 6 fail
|
fail
37%
19 pass · 32 fail
|
unsupported 51 unsupported |
fail
90%
46 pass · 5 fail
|
Kimi (K2.x section tokens, K3 XTML)
kimi |
fail
71%
35 pass · 14 fail
|
unsupported 49 unsupported |
fail
90%
44 pass · 5 fail
|
unsupported 49 unsupported |
fail
45%
22 pass · 19 soft pass · 8 fail
|
Llama (3.1, 3.2, 3.3, 4)
llama |
fail
58%
14 pass · 10 fail · 10 unsupported
|
unsupported 34 unsupported |
fail
26%
9 pass · 25 fail
|
unsupported 34 unsupported |
fail
26%
9 pass · 25 fail
|
Mistral / Magistral / Ministral / Devstral ([TOOL_CALLS])
mistral |
fail
60%
28 pass · 19 fail
|
pass 100% 14 pass · 33 unsupported synthetic-strategy failures (not counted): 2 |
fail
49%
23 pass · 1 soft pass · 23 fail
|
unsupported 47 unsupported |
fail
81%
38 pass · 2 soft pass · 7 fail
|
Qwen3 (Hermes-style JSON tool calls)
qwen3-hermes |
fail
46%
23 pass · 16 soft pass · 11 fail
|
fail
78%
39 pass · 9 soft pass · 2 fail
|
fail
14%
7 pass · 4 soft pass · 39 fail
|
unsupported 50 unsupported |
fail
46%
23 pass · 18 soft pass · 9 fail
|
Qwen XML tool calls (Qwen3-Coder, Qwen3.5/3.6/3.8)
qwen3-xml |
fail
44%
24 pass · 23 soft pass · 7 fail
|
fail
83%
45 pass · 9 fail
|
fail
35%
19 pass · 27 soft pass · 8 fail
|
unsupported 54 unsupported |
fail
44%
24 pass · 25 soft pass · 5 fail
|
Pass rate by check
Over all supported fixtures of each run. A fixture counts once per check: its worst result over the non-streaming parse and every realistic chunking strategy.
| Check | llamacpp a25c9865 | ollama 7af39318 | sglang 0.5.20 | transformers 5.17.0 | vllm 0.30.0 |
|---|---|---|---|---|---|
expected_match |
67% 273/405 (46 soft) | 87% 208/240 (10 soft) | 43% 176/414 (62 soft) | 68% 30/44 (7 soft) | 61% 264/434 (86 soft) |
expected_error |
27% 9/33 | 81% 17/21 | 53% 18/34 | 75% 3/4 | 29% 10/35 |
stream_equals_nonstream |
91% 397/438 | 98% 255/261 | 52% 233/448 (33 soft) | 81% 39/48 (7 soft) | 75% 354/469 (26 soft) |
split_invariance |
n/a | n/a | 66% 296/448 (36 soft) | n/a | 90% 421/469 (1 soft) |
no_leakage |
97% 425/438 | 98% 256/261 | 86% 387/448 | 100% 48/48 | 97% 454/469 |
arguments_json |
85% 263/308 | 100% 193/193 | 76% 241/319 | 100% 32/32 | 89% 341/384 |
arguments_schema |
85% 262/308 | 100% 193/193 | 70% 224/319 | 100% 32/32 | 87% 335/384 |
parallel_order |
74% 34/46 | 92% 24/26 | 40% 19/48 | 100% 4/4 | 92% 48/52 |
Runs
| Engine | Version | Commit | Finished | Platform | Python | canitoolcall | Fixtures digest | Strategies | Results |
|---|---|---|---|---|---|---|---|---|---|
| llamacpp | a25c9865 | a25c9865fe03 |
linux-x86_64 | 3.12.14 | 0.1.0.dev0 | a1e7b55367e1 |
one, special, token, rand:1:8, rand:2:8, rand:3:8, rand:4:8, rand:5:8 not counted: one, special, rand:1:8, rand:2:8, rand:3:8, rand:4:8, rand:5:8 |
JSON | |
| ollama | 7af39318 | 7af393188def |
linux-x86_64 | 3.12.3 | 0.1.0.dev0 | a1e7b55367e1 |
one, special, token, rand:1:8, rand:2:8, rand:3:8, rand:4:8, rand:5:8 not counted: one, special, rand:1:8, rand:2:8, rand:3:8, rand:4:8, rand:5:8 |
JSON | |
| sglang | 0.5.20 | n/a | linux-x86_64 | 3.12.14 | 0.1.0.dev0 | a1e7b55367e1 |
one, special, token, rand:1:8, rand:2:8, rand:3:8, rand:4:8, rand:5:8 | JSON | |
| transformers | 5.17.0 | n/a | linux-x86_64 | 3.12.14 | 0.1.0.dev0 | a1e7b55367e1 |
one, special, token, rand:1:8, rand:2:8, rand:3:8, rand:4:8, rand:5:8 not counted: one, special, rand:1:8, rand:2:8, rand:3:8, rand:4:8, rand:5:8 |
JSON | |
| vllm | 0.30.0 | n/a | linux-x86_64 | 3.12.14 | 0.1.0.dev0 | a1e7b55367e1 |
one, special, token, rand:1:8, rand:2:8, rand:3:8, rand:4:8, rand:5:8 | JSON |