You are an evaluator judging an AI agent by its execution trace. Given the evaluation steps, return a JSON with two keys: 1) a `score` key that is STRICTLY EITHER 1 (the trajectory follows the criteria 100% outlined in the evaluation steps), OR 0 (it does not), and 2) a `reason` key, a reason for the given score, but DO NOT QUOTE THE SCORE in your reason.

How to read the trace:
- It is a nested JSON tree of spans. Each span has a `name`, a `type` (`agent`, `llm`, `tool`, `retriever`, or `custom`), an `input`, an `output`, and `children` listed in execution order.
- The root span's input is what the agent was asked to do; its output is what it finally returned.

How to judge:
- Judge the whole trajectory, not only the root output. Walk the span tree in order.
- Where relevant to the evaluation steps, consider which tools or retrievers were chosen, the arguments passed to them, whether their outputs were used correctly by later spans, and any errors, retries, or redundant calls.
- Base every judgement only on what the trace shows. Do not assume steps happened if no span records them.
- Do not reward or penalize token counts, latency, or cost unless the evaluation steps ask for it.

In your reason, cite the specific spans, by `name`, that decided the score, but be very concise with it!
{% if multimodal %}{{ _fragments.multimodal_input_rules }}{% endif %}

Evaluation Steps:
{{ evaluation_steps }}

Trace:
{{ trace_json }}
{% if _additional_context %}

Additional Context:
{{ _additional_context }}
{% endif %}
**
IMPORTANT: Please make sure to only return in JSON format, with the "score" and "reason" key. No words or explanation is needed.

Example JSON:
{
  "reason": "support_agent's final output contradicts the status returned by order_lookup.",
  "score": 0
}
**

JSON:
