Tool Calling and Reasoning

Introduction

An LLM response can mix up to three parts: a reasoning section (e.g. <think>...</think>), plain text output, and one or more tool calls. Different models use different tags and orderings for these parts. XGrammar provides Builtin Structural Tag to generate a StructuralTag that describes the full output structure for a given model, so all three parts can be constrained during decoding.

The API accepts tools and tool_choice in the OpenAI Chat Completions convention. Serving engines already using that format can adopt XGrammar with minimal changes. For builtin / hosted tools (e.g. web_search_preview), the API extends the convention with XGrammar-specific fields; see OpenAI Tool Call Schema for type definitions.

The tool_choice parameter controls how tool calls and text are mixed:

  • Auto ("auto"): the model may output plain text, tool calls, or both.

  • Required ("required"): at least one tool call is required; plain-text-only output is not allowed.

  • Forced (named choice): exactly one specified tool must be called.

  • None ("none"): tool calls are disabled; only text and reasoning are allowed.

The reasoning parameter controls the model-specific reasoning section.

Basic API: get_model_structural_tag

get_model_structural_tag generates a StructuralTag for the given model type with the specified tools and options. The returned StructuralTag can be used with Grammar.from_structural_tag or GrammarCompiler.compile_structural_tag to obtain the corresponding grammar.

Use it when you need to constrain the model to output in a fixed pattern such as “tool name + parameter JSON”, e.g. for Llama, Qwen, Kimi, DeepSeek, OpenAI Harmony, etc.

Parameters

  • model (str): The structural-tag style. Valid values are "llama", "qwen_3", "qwen_3_5", "qwen_3_coder", "kimi", "kimi_k3", "deepseek_r1", "deepseek_v3_1", "harmony", "deepseek_v3_2", "minimax", "minimax_m3", "glm_4_7", "deepseek_v4", "cohere", "exaone".

  • tools (List[ToolParam | dict], optional): Function and builtin tools available to the model. The list can contain two kinds of tools:

    • Function tools use the OpenAI Chat Completions shape:

      {"type": "function", "function": {"name": "...", "parameters": {...}}}
      

      The "parameters" field accepts a JSON Schema dict, True (any JSON), or can be omitted (unconstrained). When "strict" is False, the parameters constraint is skipped. MiniMax M3 currently requires fixed property names and rejects unconstrained parameters and schemas with runtime-named properties.

    • Builtin tools use a compact shape with XGrammar-specific fields:

      {"type": "web_search_preview", "name": "browser.search", "parameters": {...}}
      
      • type: the provider-level builtin tool type.

      • name: the model-output tool name (defaults to type if omitted).

      • parameters: the JSON schema for constrained decoding of the builtin tool arguments.

    Default None (treated as empty list).

  • tool_choice (ToolChoiceOptionParam | dict | None, optional): Controls whether the model may or must call tools. Default "auto".

    • "auto": the model chooses between text output and tool calls.

    • None: treated the same as "auto".

    • "none": disables all tools.

    • "required": requires at least one tool call.

    • {"type": "function", "function": {"name": ...}}: forces one function tool.

    • {"type": <builtin_type>}: forces one builtin tool (matched by type).

    • {"type": "allowed_tools", "allowed_tools": {"mode": ..., "tools": [...]}}: limits available tools before applying its mode. The tools list may contain both function refs and builtin refs (matched by type).

  • reasoning ("enabled" | "disabled" | "auto" | bool, optional): Controls the model-specific reasoning section. The string modes are recommended. The boolean aliases True and False are deprecated but remain supported as "enabled" and "disabled", respectively. For models with a leading reasoning block, "auto" means that the prompt does not prefill its opener, and the model may emit one complete block or answer/call a tool directly. Default "enabled" in get_model_structural_tag; the model-specific get_minimax_m3_structural_tag defaults to "auto". For MiniMax M3, these modes must be paired with the chat template’s enabled, disabled, and adaptive thinking modes; adaptive output may start with </mm:think> when skipping reasoning. Boolean aliases are supported only by get_model_structural_tag; model-specific builders take the three explicit string modes.

  • any_order (bool, optional): When True, applies any_order=True to every JSONSchemaFormat in the generated structural tag, so each tool’s arguments may be emitted in any property order (see JSONSchemaFormat for the exact semantics). Default False, which keeps the declared property order with full validation.

  • max_whitespace_cnt (Optional[int], optional): Caps the number of consecutive whitespace characters. Setting it (e.g. 2) bounds runs of whitespace, which avoids the unbounded-whitespace outputs some models emit in bad cases that would otherwise blow up grammar compilation/matching.

Passing an unsupported model, invalid tool_choice, or invalid reasoning mode will raise ValueError.

Returns

StructuralTag: The structural tag for the given model’s function-calling format.

Examples

Function tools

from xgrammar import Grammar, get_model_structural_tag

tools = [
    {
        "type": "function",
        "function": {
            "name": "get_weather",
            "parameters": {
                "type": "object",
                "properties": {"city": {"type": "string"}},
            },
        },
    },
    {
        "type": "function",
        "function": {
            "name": "get_time",
            "parameters": {"type": "object", "properties": {}},
        },
    },
]

structural_tag = get_model_structural_tag("llama", tools=tools)
grammar = Grammar.from_structural_tag(structural_tag)

Builtin tools

For models that support builtin tools (e.g. Harmony / gpt-oss), include builtin tools in the same tools list:

structural_tag = get_model_structural_tag(
    "harmony",
    tools=[
        {
            "type": "function",
            "function": {
                "name": "user_tool",
                "parameters": {"type": "object", "properties": {"q": {"type": "string"}}},
            },
        },
        {
            "type": "web_search_preview",
            "name": "browser.search",
            "parameters": {
                "type": "object",
                "properties": {"query": {"type": "string"}},
                "required": ["query"],
            },
        },
    ],
)
grammar = Grammar.from_structural_tag(structural_tag)

Reasoning mode

For formats that support reasoning (like Qwen3, DeepSeek-R1, Kimi-K2), use the recommended string modes:

structural_tag = get_model_structural_tag("qwen_3", tools=tools, reasoning="enabled")
grammar = Grammar.from_structural_tag(structural_tag)

If reasoning is omitted, reasoning is enabled by default. Adaptive reasoning must be requested explicitly:

structural_tag = get_model_structural_tag(
    "qwen_3", tools=tools, reasoning="auto"
)

The model-specific get_minimax_m3_structural_tag builder is the exception: its direct-call default is reasoning="auto".

Tool choice

Force a specific function tool:

structural_tag = get_model_structural_tag(
    "llama",
    tools=tools,
    tool_choice={"type": "function", "function": {"name": "get_weather"}},
)

Force a builtin tool by type:

structural_tag = get_model_structural_tag(
    "harmony",
    tools=[...],
    tool_choice={"type": "web_search_preview"},
)

Allow only a subset of tools:

structural_tag = get_model_structural_tag(
    "harmony",
    tools=[...],
    tool_choice={
        "type": "allowed_tools",
        "allowed_tools": {
            "mode": "auto",
            "tools": [
                {"type": "function", "function": {"name": "get_weather"}},
                {"type": "web_search_preview"},
            ],
        },
    },
)

Special tokens in text

By default, special tokens like <think> / </think> are not allowed in free text. Pass exclude_special_tokens=False to allow them:

structural_tag = get_model_structural_tag(
    "qwen_3",
    tools=tools,
    exclude_special_tokens=False,
)

Defaults to True. No effect for models without special tokens (e.g. harmony).


Supported models

The model argument of get_model_structural_tag accepts the style names below:

model (style)

Supported models

"llama"

Meta-Llama-3, Llama-3.1, Llama-3.2

"qwen_3"

Qwen3, Qwen3-Next

"qwen_3_5"

Qwen3.5, Qwen3.6

"qwen_3_coder"

Qwen3-Coder, Qwen3-Coder-Next

"kimi"

Kimi-K2, Kimi-K2.5

"kimi_k3"

Kimi-K3

"deepseek_r1"

DeepSeek-R1, DeepSeek-R1-0528

"deepseek_v3_1"

DeepSeek-V3.1, DeepSeek-V3.2-Exp

"harmony"

gpt-oss

"deepseek_v3_2"

DeepSeek-V3.2

"minimax"

MiniMax-M2.5

"minimax_m3"

MiniMax-M3

"glm_4_7"

GLM-5, GLM-4.7

"deepseek_v4"

DeepSeek-V4

"cohere"

Cohere Command models using XML tool calls

"exaone"

EXAONE-4.0-32B, EXAONE-4.0-1.2B

Extending with custom models

Use register_model_structural_tag to add support for a new model format. See the Builtin Structural Tag API Reference for details.

Next Steps