CanIToolCall caniuse.com for tool calling

Tool-call parser conformance matrix

Each cell shows how one inference engine's own tool-call and reasoning parsers handle one model family's recorded outputs, parsed once without streaming and once per token-chunking strategy. Pass rates are strict; whitespace-only differences are counted separately as soft passes. Latest run: .

Families × engines

Headline status, strict pass rate and case counts per family and engine version. Checks that did not pass on every case are listed in each cell; select a cell for the failing fixtures.
Family llamacpp a25c9865 run ollama 7af39318 run sglang 0.5.20 run transformers 5.17.0 run vllm 0.30.0 run
DeepSeek (V3/R1, V3.1, V3.2 DSML, V4 DSML, V4.1 DSML) deepseek fail 55% 35 pass · 29 fail · 21 unsupported
  • expected_match 33/58
  • expected_error 2/6
  • stream_equals_nonstream 61/64
  • arguments_json 24/28
  • arguments_schema 24/28
  • parallel_order 4/7
synthetic-strategy failures (not counted): 29
pass 100% 12 pass · 73 unsupported fail 20% 13 pass · 23 soft pass · 28 fail · 21 unsupported
  • expected_match 11/58 (23 soft)
  • expected_error 2/6
  • stream_equals_nonstream 13/64 (25 soft)
  • split_invariance 13/64 (29 soft)
  • no_leakage 60/64
  • arguments_json 47/53
  • arguments_schema 33/53
  • parallel_order 4/7
unsupported 85 unsupported fail 51% 43 pass · 9 soft pass · 33 fail
  • expected_match 43/78 (9 soft)
  • expected_error 1/7
  • stream_equals_nonstream 56/85
  • split_invariance 60/85
  • no_leakage 83/85
  • arguments_json 54/72
  • arguments_schema 50/72
  • parallel_order 8/11
Gemma 4 (<|tool_call>call:NAME{...} object notation) gemma4 fail 73% 35 pass · 5 soft pass · 8 fail
  • expected_match 35/44 (5 soft)
  • expected_error 2/4
  • stream_equals_nonstream 46/48
  • arguments_json 36/39
  • arguments_schema 36/39
synthetic-strategy failures (not counted): 8
fail 83% 40 pass · 1 soft pass · 7 fail
  • expected_match 37/44 (1 soft)
  • expected_error 3/4
  • no_leakage 47/48
synthetic-strategy failures (not counted): 7
fail 62% 30 pass · 5 soft pass · 13 fail
  • expected_match 30/44 (5 soft)
  • expected_error 2/4
  • stream_equals_nonstream 38/48
  • arguments_json 35/39
  • arguments_schema 33/39
  • parallel_order 3/4
fail 69% 33 pass · 7 soft pass · 8 fail
  • expected_match 30/44 (7 soft)
  • expected_error 3/4
  • stream_equals_nonstream 39/48 (7 soft)
synthetic-strategy failures (not counted): 8
fail 62% 30 pass · 2 soft pass · 16 fail
  • expected_match 30/44 (2 soft)
  • expected_error 0/4
  • stream_equals_nonstream 42/48 (2 soft)
  • no_leakage 47/48
  • arguments_schema 39/41
GLM (4.5, 4.6, 4.7, 5.x) glm fail 80% 41 pass · 2 soft pass · 8 fail
  • expected_match 41/48 (2 soft)
  • expected_error 0/3
  • stream_equals_nonstream 48/51
  • no_leakage 50/51
  • arguments_json 37/43
  • arguments_schema 37/43
  • parallel_order 7/8
synthetic-strategy failures (not counted): 8
fail 91% 29 pass · 3 fail · 19 unsupported
  • expected_match 28/30
  • expected_error 1/2
  • no_leakage 31/32
synthetic-strategy failures (not counted): 3
fail 49% 25 pass · 2 soft pass · 24 fail
  • expected_match 25/48 (2 soft)
  • expected_error 0/3
  • stream_equals_nonstream 26/51 (2 soft)
  • split_invariance 32/51 (7 soft)
  • no_leakage 39/51
  • arguments_json 38/44
  • arguments_schema 38/44
  • parallel_order 0/8
unsupported 51 unsupported fail 65% 33 pass · 11 soft pass · 7 fail
  • expected_match 33/48 (11 soft)
  • expected_error 0/3
  • stream_equals_nonstream 35/51 (11 soft)
  • split_invariance 50/51 (1 soft)
  • no_leakage 49/51
  • arguments_json 41/45
  • arguments_schema 41/45
gpt-oss (Harmony response format) gpt-oss fail 84% 43 pass · 8 fail
  • expected_match 42/45
  • expected_error 1/6
  • stream_equals_nonstream 49/51
  • arguments_json 35/36
  • arguments_schema 35/36
  • parallel_order 0/1
synthetic-strategy failures (not counted): 8
fail 88% 45 pass · 6 fail
  • expected_match 41/45
  • expected_error 5/6
  • stream_equals_nonstream 47/51
  • parallel_order 0/1
synthetic-strategy failures (not counted): 6
fail 37% 19 pass · 32 fail
  • expected_match 13/45
  • stream_equals_nonstream 21/51
  • split_invariance 41/51
  • no_leakage 22/51
  • parallel_order 0/1
unsupported 51 unsupported fail 90% 46 pass · 5 fail
  • expected_match 44/45
  • expected_error 3/6
  • stream_equals_nonstream 50/51
  • no_leakage 49/51
  • arguments_json 35/37
  • arguments_schema 35/37
Kimi (K2.x section tokens, K3 XTML) kimi fail 71% 35 pass · 14 fail
  • expected_match 35/46
  • expected_error 1/3
  • stream_equals_nonstream 39/49
  • no_leakage 48/49
  • arguments_json 26/31
  • arguments_schema 25/31
  • parallel_order 3/4
synthetic-strategy failures (not counted): 14
unsupported 49 unsupported fail 90% 44 pass · 5 fail
  • expected_match 42/46
  • expected_error 2/3
  • stream_equals_nonstream 45/49
  • arguments_json 34/36
  • arguments_schema 34/36
unsupported 49 unsupported fail 45% 22 pass · 19 soft pass · 8 fail
  • expected_match 20/46 (19 soft)
  • expected_error 2/3
  • stream_equals_nonstream 43/49
  • split_invariance 48/49
  • no_leakage 47/49
  • arguments_json 34/36
  • arguments_schema 34/36
Llama (3.1, 3.2, 3.3, 4) llama fail 58% 14 pass · 10 fail · 10 unsupported
  • expected_match 14/23
  • expected_error 0/1
  • stream_equals_nonstream 20/24
  • arguments_json 14/19
  • arguments_schema 14/19
synthetic-strategy failures (not counted): 10
unsupported 34 unsupported fail 26% 9 pass · 25 fail
  • expected_match 9/32
  • expected_error 1/2
  • stream_equals_nonstream 13/34
  • split_invariance 15/34
  • no_leakage 33/34
  • arguments_json 7/26
  • arguments_schema 7/26
unsupported 34 unsupported fail 26% 9 pass · 25 fail
  • expected_match 9/32
  • expected_error 0/2
  • stream_equals_nonstream 9/34
  • split_invariance 16/34
  • arguments_json 25/28
  • arguments_schema 25/28
Mistral / Magistral / Ministral / Devstral ([TOOL_CALLS]) mistral fail 60% 28 pass · 19 fail
  • expected_match 28/45
  • expected_error 0/2
  • stream_equals_nonstream 36/47
  • no_leakage 45/47
  • arguments_json 21/33
  • arguments_schema 21/33
  • parallel_order 3/8
synthetic-strategy failures (not counted): 19
pass 100% 14 pass · 33 unsupported synthetic-strategy failures (not counted): 2 fail 49% 23 pass · 1 soft pass · 23 fail
  • expected_match 21/45 (1 soft)
  • stream_equals_nonstream 26/47
  • split_invariance 27/47
  • no_leakage 42/47
  • arguments_json 27/34
  • arguments_schema 27/34
  • parallel_order 0/8
unsupported 47 unsupported fail 81% 38 pass · 2 soft pass · 7 fail
  • expected_match 38/45 (2 soft)
  • expected_error 1/2
  • stream_equals_nonstream 41/47
  • split_invariance 44/47
  • no_leakage 43/47
  • arguments_json 33/39
  • arguments_schema 33/39
  • parallel_order 7/8
Qwen3 (Hermes-style JSON tool calls) qwen3-hermes fail 46% 23 pass · 16 soft pass · 11 fail
  • expected_match 21/43 (16 soft)
  • expected_error 3/7
  • stream_equals_nonstream 49/50
  • no_leakage 45/50
  • arguments_json 28/32
  • arguments_schema 28/32
  • parallel_order 6/7
synthetic-strategy failures (not counted): 11
fail 78% 39 pass · 9 soft pass · 2 fail
  • expected_match 33/43 (9 soft)
  • expected_error 6/7
synthetic-strategy failures (not counted): 2
fail 14% 7 pass · 4 soft pass · 39 fail
  • expected_match 6/43 (4 soft)
  • expected_error 3/7
  • stream_equals_nonstream 11/50
  • split_invariance 17/50
  • no_leakage 44/50
  • arguments_json 0/32
  • arguments_schema 0/32
  • parallel_order 0/7
unsupported 50 unsupported fail 46% 23 pass · 18 soft pass · 9 fail
  • expected_match 23/43 (18 soft)
  • expected_error 3/7
  • stream_equals_nonstream 33/50 (8 soft)
  • no_leakage 48/50
  • arguments_json 32/37
  • arguments_schema 32/37
Qwen XML tool calls (Qwen3-Coder, Qwen3.5/3.6/3.8) qwen3-xml fail 44% 24 pass · 23 soft pass · 7 fail
  • expected_match 24/53 (23 soft)
  • expected_error 0/1
  • stream_equals_nonstream 49/54
  • no_leakage 50/54
  • arguments_json 42/47
  • arguments_schema 42/47
synthetic-strategy failures (not counted): 7
fail 83% 45 pass · 9 fail
  • expected_match 44/53
  • stream_equals_nonstream 52/54
  • no_leakage 51/54
  • parallel_order 6/7
synthetic-strategy failures (not counted): 10
fail 35% 19 pass · 27 soft pass · 8 fail
  • expected_match 19/53 (27 soft)
  • expected_error 0/1
  • stream_equals_nonstream 40/54 (6 soft)
  • no_leakage 50/54
  • arguments_json 47/49
  • arguments_schema 46/49
  • parallel_order 6/7
unsupported 54 unsupported fail 44% 24 pass · 25 soft pass · 5 fail
  • expected_match 24/53 (25 soft)
  • expected_error 0/1
  • stream_equals_nonstream 45/54 (5 soft)
  • arguments_json 46/49
  • arguments_schema 46/49

Pass rate by check

Over all supported fixtures of each run. A fixture counts once per check: its worst result over the non-streaming parse and every realistic chunking strategy.

Strict pass rate per check and engine version (strict passes / fixtures the check applied to; soft passes in parentheses).
Check llamacpp a25c9865ollama 7af39318sglang 0.5.20transformers 5.17.0vllm 0.30.0
expected_match 67% 273/405 (46 soft) 87% 208/240 (10 soft) 43% 176/414 (62 soft) 68% 30/44 (7 soft) 61% 264/434 (86 soft)
expected_error 27% 9/33 81% 17/21 53% 18/34 75% 3/4 29% 10/35
stream_equals_nonstream 91% 397/438 98% 255/261 52% 233/448 (33 soft) 81% 39/48 (7 soft) 75% 354/469 (26 soft)
split_invariance n/a n/a 66% 296/448 (36 soft) n/a 90% 421/469 (1 soft)
no_leakage 97% 425/438 98% 256/261 86% 387/448 100% 48/48 97% 454/469
arguments_json 85% 263/308 100% 193/193 76% 241/319 100% 32/32 89% 341/384
arguments_schema 85% 262/308 100% 193/193 70% 224/319 100% 32/32 87% 335/384
parallel_order 74% 34/46 92% 24/26 40% 19/48 100% 4/4 92% 48/52

Runs

The run behind each column. When an engine version was run more than once, the latest run is shown.
EngineVersionCommit FinishedPlatformPython canitoolcallFixtures digestStrategies Results
llamacpp a25c9865 a25c9865fe03 linux-x86_64 3.12.14 0.1.0.dev0 a1e7b55367e1 one, special, token, rand:1:8, rand:2:8, rand:3:8, rand:4:8, rand:5:8
not counted: one, special, rand:1:8, rand:2:8, rand:3:8, rand:4:8, rand:5:8
JSON
ollama 7af39318 7af393188def linux-x86_64 3.12.3 0.1.0.dev0 a1e7b55367e1 one, special, token, rand:1:8, rand:2:8, rand:3:8, rand:4:8, rand:5:8
not counted: one, special, rand:1:8, rand:2:8, rand:3:8, rand:4:8, rand:5:8
JSON
sglang 0.5.20 n/a linux-x86_64 3.12.14 0.1.0.dev0 a1e7b55367e1 one, special, token, rand:1:8, rand:2:8, rand:3:8, rand:4:8, rand:5:8 JSON
transformers 5.17.0 n/a linux-x86_64 3.12.14 0.1.0.dev0 a1e7b55367e1 one, special, token, rand:1:8, rand:2:8, rand:3:8, rand:4:8, rand:5:8
not counted: one, special, rand:1:8, rand:2:8, rand:3:8, rand:4:8, rand:5:8
JSON
vllm 0.30.0 n/a linux-x86_64 3.12.14 0.1.0.dev0 a1e7b55367e1 one, special, token, rand:1:8, rand:2:8, rand:3:8, rand:4:8, rand:5:8 JSON