Tool Calling and Reasoning

Introduction

An LLM response can mix up to three parts: a reasoning section (e.g. <think>...</think>), plain text output, and one or more tool calls. Different models use different tags and orderings for these parts. XGrammar provides Builtin Structural Tag to generate a StructuralTag that describes the full output structure for a given model, so all three parts can be constrained during decoding.

The API accepts tools and tool_choice in the OpenAI Chat Completions convention. Serving engines already using that format can adopt XGrammar with minimal changes. For builtin / hosted tools (e.g. web_search_preview), the API extends the convention with XGrammar-specific fields; see OpenAI Tool Call Schema for type definitions.

The tool_choice parameter controls how tool calls and text are mixed:

  • Auto ("auto"): the model may output plain text, tool calls, or both.

  • Required ("required"): at least one tool call is required; plain-text-only output is not allowed.

  • Forced (named choice): exactly one specified tool must be called.

  • None ("none"): tool calls are disabled; only text and reasoning are allowed.

The reasoning parameter controls the model-specific reasoning section.

Basic API: get_model_structural_tag

get_model_structural_tag generates a StructuralTag for the given model type with the specified tools and options. The returned StructuralTag can be used with Grammar.from_structural_tag or GrammarCompiler.compile_structural_tag to obtain the corresponding grammar.

Use it when you need to constrain the model to output in a fixed pattern such as “tool name + parameter JSON”, e.g. for Llama, Qwen, Kimi, DeepSeek, OpenAI Harmony, etc.

Parameters

  • model (str): The structural-tag style. Valid values are "llama", "qwen_3", "qwen_3_5", "qwen_3_coder", "kimi", "kimi_k3", "deepseek_r1", "deepseek_v3_1", "harmony", "deepseek_v3_2", "minimax", "minimax_m3", "glm_4_7", "deepseek_v4", "deepseek_v4_1", "mimo", "cohere", "exaone", "gemma_4".

  • tools (List[ToolParam | dict], optional): Function and builtin tools available to the model. The list can contain two kinds of tools:

    • Function tools use the OpenAI Chat Completions shape:

      {"type": "function", "function": {"name": "...", "parameters": {...}}}
      

      The "parameters" field accepts a JSON Schema dict, True (any JSON), or can be omitted (unconstrained). When "strict" is False, the parameters constraint is skipped. MiniMax M3 currently requires fixed property names and rejects unconstrained parameters and schemas with runtime-named properties.

    • Builtin tools use a compact shape with XGrammar-specific fields:

      {"type": "web_search_preview", "name": "browser.search", "parameters": {...}}
      
      • type: the provider-level builtin tool type.

      • name: the model-output tool name (defaults to type if omitted).

      • parameters: the JSON schema for constrained decoding of the builtin tool arguments.

    Default None (treated as empty list).

  • tool_choice (ToolChoiceOptionParam | dict | None, optional): Controls whether the model may or must call tools. Default "auto".

    • "auto": the model chooses between text output and tool calls.

    • None: treated the same as "auto".

    • "none": disables all tools.

    • "required": requires at least one tool call.

    • {"type": "function", "function": {"name": ...}}: forces one function tool.

    • {"type": <builtin_type>}: forces one builtin tool (matched by type).

    • {"type": "allowed_tools", "allowed_tools": {"mode": ..., "tools": [...]}}: limits available tools before applying its mode. The tools list may contain both function refs and builtin refs (matched by type).

  • reasoning ("enabled" | "disabled" | "auto" | bool, optional): Controls the model-specific reasoning section. The string modes are recommended. The boolean aliases True and False are deprecated but remain supported as "enabled" and "disabled", respectively. For models with a leading reasoning block, "auto" means that the prompt does not prefill its opener, and the model may emit one complete block or answer/call a tool directly. Default "enabled" in get_model_structural_tag; the model-specific get_minimax_m3_structural_tag defaults to "auto". For MiniMax M3, these modes must be paired with the chat template’s enabled, disabled, and adaptive thinking modes; adaptive output may start with </mm:think> when skipping reasoning. Boolean aliases are supported only by get_model_structural_tag; model-specific builders take the three explicit string modes.

  • any_order (bool, optional): When True, applies any_order=True to every JSONSchemaFormat in the generated structural tag, so each tool’s arguments may be emitted in any property order (see JSONSchemaFormat for the exact semantics). Default False, which keeps the declared property order with full validation.

  • parallel_tool_calls (bool, optional): Whether the model may emit more than one tool call in a single response, following the OpenAI Chat Completions parameter of the same name. Default True, which keeps the multi-call grammar. When False, the structural tag allows at most one tool call: exactly zero or one under tool_choice="auto", exactly one under "required" or a forced tool. Free text is still allowed before the call, but generation must end once the call is closed, since a second call is no longer reachable.

  • token_markers (bool, optional): Match control markers (<tool_call>, </think>, …) as their dedicated tokens instead of as strings; see Token-level markers. Supported for glm_4_7, qwen_3, qwen_3_5 and qwen_3_coder. Default False.

  • max_whitespace_cnt (Optional[int], optional): Caps the number of consecutive whitespace characters. Setting it (e.g. 2) bounds runs of whitespace, which avoids the unbounded-whitespace outputs some models emit in bad cases that would otherwise blow up grammar compilation/matching.

Passing an unsupported model, invalid tool_choice, or invalid reasoning mode will raise ValueError.

Returns

StructuralTag: The structural tag for the given model’s function-calling format.

Examples

Function tools

from xgrammar import Grammar, get_model_structural_tag

tools = [
    {
        "type": "function",
        "function": {
            "name": "get_weather",
            "parameters": {
                "type": "object",
                "properties": {"city": {"type": "string"}},
            },
        },
    },
    {
        "type": "function",
        "function": {
            "name": "get_time",
            "parameters": {"type": "object", "properties": {}},
        },
    },
]

structural_tag = get_model_structural_tag("llama", tools=tools)
grammar = Grammar.from_structural_tag(structural_tag)

Builtin tools

For models that support builtin tools (e.g. Harmony / gpt-oss), include builtin tools in the same tools list:

structural_tag = get_model_structural_tag(
    "harmony",
    tools=[
        {
            "type": "function",
            "function": {
                "name": "user_tool",
                "parameters": {"type": "object", "properties": {"q": {"type": "string"}}},
            },
        },
        {
            "type": "web_search_preview",
            "name": "browser.search",
            "parameters": {
                "type": "object",
                "properties": {"query": {"type": "string"}},
                "required": ["query"],
            },
        },
    ],
)
grammar = Grammar.from_structural_tag(structural_tag)

Reasoning mode

For formats that support reasoning (like Qwen3, DeepSeek-R1, Kimi-K2), use the recommended string modes:

structural_tag = get_model_structural_tag("qwen_3", tools=tools, reasoning="enabled")
grammar = Grammar.from_structural_tag(structural_tag)

If reasoning is omitted, reasoning is enabled by default. Adaptive reasoning must be requested explicitly:

structural_tag = get_model_structural_tag(
    "qwen_3", tools=tools, reasoning="auto"
)

The model-specific get_minimax_m3_structural_tag builder is the exception: its direct-call default is reasoning="auto".

Tool choice

Force a specific function tool:

structural_tag = get_model_structural_tag(
    "llama",
    tools=tools,
    tool_choice={"type": "function", "function": {"name": "get_weather"}},
)

Force a builtin tool by type:

structural_tag = get_model_structural_tag(
    "harmony",
    tools=[...],
    tool_choice={"type": "web_search_preview"},
)

Allow only a subset of tools:

structural_tag = get_model_structural_tag(
    "harmony",
    tools=[...],
    tool_choice={
        "type": "allowed_tools",
        "allowed_tools": {
            "mode": "auto",
            "tools": [
                {"type": "function", "function": {"name": "get_weather"}},
                {"type": "web_search_preview"},
            ],
        },
    },
)

Special tokens in text

By default, special tokens like <think> / </think> are not allowed in free text. Pass exclude_special_tokens=False to allow them:

structural_tag = get_model_structural_tag(
    "qwen_3",
    tools=tools,
    exclude_special_tokens=False,
)

Defaults to True. No effect for models without special tokens (e.g. harmony).

Parallel tool calls

By default the grammar accepts as many tool calls as the model wants to emit. Pass parallel_tool_calls=False to cap the response at a single call:

structural_tag = get_model_structural_tag(
    "qwen_3",
    tools=tools,
    parallel_tool_calls=False,
)

With tool_choice="auto" the model may still answer with plain text instead, so the response holds exactly zero or one call; with "required" or a forced tool it holds exactly one. A forced tool choice already resolves to a single call, so the flag does not change it.

Token-level markers

By default control markers such as <tool_call> are strings, so the grammar also accepts them spelled from ordinary sub-tokens (<, tool_call, >), which a parser that recognises markers by token ID cannot parse. token_markers=True requires the dedicated tokens instead; the tokenizer must define each marker as a single token, and the tag must be compiled with a tokenizer-configured GrammarCompiler, not Grammar.from_structural_tag. Markers in token-triggered free text are then also only excluded as tokens, so use it only with such parsers. Markers inside schema-driven content (e.g. GLM’s <arg_key>) stay strings.


Supported models

The model argument of get_model_structural_tag accepts the style names below:

model (style)

Supported models

"llama"

Meta-Llama-3, Llama-3.1, Llama-3.2

"qwen_3"

Qwen3, Qwen3-Next

"qwen_3_5"

Qwen3.5, Qwen3.6

"qwen_3_coder"

Qwen3-Coder, Qwen3-Coder-Next

"kimi"

Kimi-K2, Kimi-K2.5

"kimi_k3"

Kimi-K3

"deepseek_r1"

DeepSeek-R1, DeepSeek-R1-0528

"deepseek_v3_1"

DeepSeek-V3.1, DeepSeek-V3.2-Exp

"harmony"

gpt-oss

"deepseek_v3_2"

DeepSeek-V3.2

"minimax"

MiniMax-M2.5

"minimax_m3"

MiniMax-M3

"glm_4_7"

GLM-5, GLM-4.7

"deepseek_v4"

DeepSeek-V4

"deepseek_v4_1"

DeepSeek-V4.1-Flash

"mimo"

MiMo-V2.6-Pro-RL, MiMo-V2.6-Flash-RL

"cohere"

Cohere Command models using XML tool calls

"exaone"

EXAONE-4.0-32B, EXAONE-4.0-1.2B

"gemma_4"

Gemma 4 (gemma-4-E2B-it, gemma-4-12b-it, gemma-4-26b-a4b-it, gemma-4-31b-it); arguments use the "gemma" JSON-schema style

Extending with custom models

Use register_model_structural_tag to add support for a new model format. See the Builtin Structural Tag API Reference for details.

Next Steps