Extract, normalize, and repair heterogeneous tool calls from large language model text streams. Sub-0.05ms deterministic execution across standard OpenAI, Claude XML, DeepSeek DSML, Qwen, and Llama formats.
pip install agent-tool-parser
Evaluate how heterogeneous or truncated model completions are normalized to canonical ToolCall structures.
Client-side agent runtimes encounter structural inconsistencies when interfacing with diverse cloud APIs and local inference engines.
Inference endpoints (such as OpenRouter, DeepSeek API, or vLLM) frequently return empty tool_calls arrays while embedding raw XML, DSML, or proprietary delimiter tokens directly within message.content.
When model generation terminates at token limits or network interruptions occur, standard JSON deserialization fails. The parser evaluates bracket stacks to close and recover partial JSON payloads deterministically.
Serializing shell scripts, unified diffs, or Python functions within JSON string literals frequently results in escape violations. The parser extracts unescaped CDATA blocks and multi-line script bodies without string corruption.
Reference implementations for integrating agent-tool-parser into agent loops, gateway proxies, and parallel execution pipelines.
from agent_tool_parser import try_parse_tool_calls
def run_agent_turn(messages: list[dict], client, model_name: str):
response = client.chat.completions.create(
model=model_name,
messages=messages,
temperature=0.0,
)
raw_content = response.choices[0].message.content or ""
# Resilient extraction across DeepSeek, Claude, Qwen, Llama & non-standard JSON
tool_calls = try_parse_tool_calls(
raw_content,
allowed_tools={"read_file", "write_file", "bash", "search"},
tool_aliases={"view": "read_file", "sh": "bash", "cmd": "bash"}
)
if not tool_calls:
return {"action": "reply", "content": raw_content}
# Execute all tools (in parallel or sequentially)
results = []
for call in tool_calls:
output = execute_tool(call.name, call.args)
results.append({"tool": call.name, "output": output})
return {"action": "continue", "results": results}
from agent_tool_parser import parse_tool_calls, safe_json_loads, ToolError
def normalize_response(response_message):
"""Normalizes response structures when inference endpoints emit markup in message content."""
# 1. Return native tool_calls if already structured by provider SDK
if getattr(response_message, "tool_calls", None):
return [
{"name": tc.function.name, "args": safe_json_loads(tc.function.arguments)}
for tc in response_message.tool_calls
]
# 2. Extract structured invocations from text content stream
content = getattr(response_message, "content", "") or ""
try:
calls = parse_tool_calls(content)
return [{"name": c.name, "args": c.args} for c in calls]
except ToolError:
return []
import asyncio
from agent_tool_parser import parse_tool_calls
async def handle_llm_turn(model_output: str):
# Parse multiple parallel tool calls from one generation step
calls = parse_tool_calls(model_output)
# Execute all tools concurrently in parallel
tasks = [async_run_tool(call.name, call.args) for call in calls]
tool_results = await asyncio.gather(*tasks, return_exceptions=True)
return list(zip([c.name for c in calls], tool_results))
Objective comparison of common parsing, repair, and extraction approaches in AI agent runtimes.
| Evaluation Criterion | agent-tool-parser | json-repair | LangChain Core | vLLM Server-Side |
|---|---|---|---|---|
| Runtime Dependencies | 0 (Pure Python) | 0 | > 30 packages | Torch, CUDA (GBs) |
| Average Latency | < 0.05 ms | ~ 0.15 ms | ~ 1.20 ms | N/A (Server-side) |
| Reasoning Stripping (<think>) | ✓ Automatic | ✗ Crashes | Partial | ✗ No |
| DeepSeek DSML & Native Tokens | ✓ Full Support | ✗ No | ✗ No | ✓ Yes |
| Anthropic Claude XML + CDATA | ✓ Full Support | ✗ No | Partial | ✗ No |
| Llama 3.1 Python AST Expressions | ✓ Safe AST (No eval) | ✗ No | ✗ No | ✗ No |
| Truncated Streaming Tail Repair | ✓ Bracket-Stack Auto-close | ✓ Yes | ✗ Crashes | ✗ No |
| Multi-Tool Calling | ✓ Any syntax | ✗ No | JSON only | ✗ No |
System trade-offs across execution boundaries, sampler integration, and dependency stacks.
| Architectural Dimension | Client-Side Post-Parsing | Logit-Level Constrained Decoding | Secondary Model Pass |
|---|---|---|---|
| Execution Boundary | Client Runtime | Inference Engine Sampler | Secondary Endpoint / Worker |
| Inference Engine Access | Agnostic (HTTP/REST text streams) | Requires direct logit & sampler access | Requires secondary model call |
| Runtime Dependencies | Python Standard Library | PyTorch, CUDA, Triton | PyTorch, Transformers / API |
| Provider API Compatibility | ✓ Any provider returning text | ✗ Incompatible with third-party APIs | ✓ Via extra API round-trip |
| Grammar Enforcement | Deterministic text parsing & repair | FSM logit masking during token sampling | Probabilistic prompting pass |
| Format Scope | XML, DSML, Markdown JSON, AST, CDATA | Strict JSON / EBNF Grammars | Model training distribution |
Built for the reality of multi-agent frameworks, high-throughput engines, and fleet mission controls.
Mission Control • Paperclip AI • Dify
When managing fleets of background agents, a single unparsed tool call causes worker crashes, poisoned queues, and blown token budgets. agent-tool-parser acts as an ingress sanitization layer before tasks hit approval gates.
CrewAI • AutoGen (AG2) • LangGraph • PydanticAI
Local models (Ollama/vLLM) frequently emit <think> tokens or parameter aliases that trigger fatal ValidationErrors in CrewAI or PydanticAI. Our interceptor normalizes them smoothly.
SGLang • vLLM • Ollama • llama.cpp
Strict constrained decoding (XGrammar / outlines) degrades reasoning quality in models like DeepSeek R1. With agent-tool-parser, keep token generation unconstrained while extracting tool calls in <0.05ms client-side.
LiteLLM • Portkey • OpenRouter
Multi-model proxies often fail to extract vendor-specific XML or DSML out of message.content into standard tool_calls arrays. Our parser acts as a zero-dependency normalization hook.
Aider • Cline • OpenHands (OpenDevin)
Multi-line Python scripts and Bash commands constantly break standard JSON escaping. Full support for XML tags, CDATA unwrapping, and code blocks ensures lossless code execution.
Smolagents • Custom ReAct Loops
Hugging Face built CodeAgent because JSON tool-calling was notoriously fragile. agent-tool-parser gives developers the reliability of code agents without requiring sandboxed Python execution environments.
Simple, predictable functional interfaces with configurable class extensions.
Parses an LLM response string and returns all structured ToolCall objects found. Raises ToolError if no valid tool calls could be extracted.
calls = parse_tool_calls(raw_text)
for call in calls:
print(call.name, call.args)
Safe non-raising variant. Returns an empty list [] if the model response only contained conversational text without any tool calls.
calls = try_parse_tool_calls(raw_text)
if not calls:
return "No tool needed, answer directly"
Reusable, configurable parser instance. Enforces tool whitelists, normalizes custom aliases (e.g. view -> read_file), and handles argument mapping.
parser = ToolParser(
allowed_tools=["read_file", "bash"],
tool_aliases={"terminal": "bash"}
)
call = parser.parse(raw_text)
Stand-alone resilient JSON loader. Auto-repairs single quotes, trailing commas, and converts Python literals (True/False/None) without throwing exceptions.
from agent_tool_parser import safe_json_loads
data = safe_json_loads("{'path': 'a.py', 'overwrite': True,}")
# {'path': 'a.py', 'overwrite': True}