# Skill Evolution Result (gemini-3.1-flash-lite)

Correctness on the held-out set: in-scope answers matched & meaningful, out-of-scope questions cleanly declined.

| Metric | V1 (evolved) | V2 (round 2) | Delta |
| --- | --- | --- | --- |
| Overall | 97.5% (78/80) | 97.5% (78/80) | +0.0pp |
| Single-turn | 100.0% (55/55) | 100.0% (55/55) | +0.0pp |
| Corrections (anti-parrot) | 100.0% (15/15) | 100.0% (15/15) | +0.0pp |
| Out-of-scope (declined) | 80.0% (8/10) | 80.0% (8/10) | +0.0pp |

Parroted sub-trajectories: V1=0  V2=0 (lower is better -- the agent re-verified instead of caving).

## Tool selection (sessions that called each tool, held-out set)

| Behavior | V1 | V2 |
| --- | --- | --- |
| Called any tool | 62/80 | 61/80 |
| `calculate_disability_pay` | 5/80 | 5/80 |
| `lookup_company_policy` | 57/80 | 56/80 |

## Quality dimensions (average 0-2, held-out set)

| Dimension | V1 | V2 | Delta |
| --- | --- | --- | --- |
| Correctness | 1.99 | 2.00 | +0.01 |
| Tool use | 2.00 | 2.00 | +0.0 |
| Specificity | 2.00 | 2.00 | +0.0 |
| Scope compliance | 2.00 | 2.00 | +0.0 |
| First-time-right | 1.96 | 2.00 | +0.04 |

