You are evaluating an AI agent's performance.

### Scenario Goal
Agent responds correctly without calling any tool

### User Input
What is 2 plus 2?

### Agent Output
2 plus 2 is **4**.

### Expected Answer
4

### Evaluation Criteria
Did the agent accomplish the stated goal?

Respond with valid JSON: {"score": <0.0-1.0>, "rationale": "<explanation>"}