Evaluation report — replies

Customer
bookshop
Suite
replies
This run
dev — 2026-09-23T07:19:02.463640+00:00
Reference
dev — 2026-09-23T07:14:08.579107+00:00
Code version
84e0070bb599b6189a5861a382b4870c959df62b
Reference approved
2026-09-23T07:15:28.850547+00:00

Did it get worse? Yes

1 check got worse compared with the reference. 1 check is on the line: the band it measured covers its threshold. Every case could be judged. No case is suspended. The suite is unchanged from the reference. The files under test are the same as the reference. The system under test answered under the same configuration as the reference.

What was under test

The files under test are the same as the reference.

How it was set up

Configured as in the reference.

ParameterThis runReference
max_tokens200200
modelclaude-haiku-4-5claude-haiku-4-5
provideranthropicanthropic
resolved_modelclaude-haiku-4-5-20251001claude-haiku-4-5-20251001
What got worse (1)
CaseCheckWhat happenedWhy
gift-wrapllm_rubricScore fell from 0.650000 to 0.472222 — beyond the noise of this check (0.550000–0.766667 across 3 samples).mean of 3 samples (0.333333, 0.700000, 0.383333)
What could not be judged (0)

Nothing in this section.

What is set aside (0)

Nothing in this section.

What was added or removed (0)

Nothing in this section.

What got better (0)

Nothing in this section.

What stayed the same (9)
CaseCheckWhat happenedWhy
gift-wraplevenshteinScore unchanged at 0.183702.mean of 3 samples (0.191257, 0.193182, 0.166667)
opening-hourslevenshteinScore unchanged at 0.197593.mean of 3 samples (0.193277, 0.210526, 0.188976)
opening-hoursllm_rubricScore unchanged at 0.394444.mean of 3 samples (0.300000, 0.350000, 0.533333)
order-a-booklevenshteinScore unchanged at 0.273279.mean of 3 samples (0.287671, 0.264840, 0.267327)
order-a-bookllm_rubricScore unchanged at 0.538889.mean of 3 samples (0.466667, 0.533333, 0.616667)
price-checklevenshteinScore unchanged at 0.179674.mean of 3 samples (0.199134, 0.182510, 0.157377)
price-checkllm_rubricScore unchanged at 0.300000.mean of 3 samples (0.300000, 0.400000, 0.200000)
signed-copieslevenshteinScore unchanged at 0.207743.mean of 3 samples (0.196078, 0.197026, 0.230126)
signed-copiesllm_rubricScore unchanged at 0.500000.mean of 3 samples (0.433333, 0.533333, 0.533333)
Where the judge placed the calibration answers (1)
CaseCheckResultDeclared bandWhy
calibration-half-rightllm_rubric0.838889 inside0.700000–0.950000mean of 3 samples (0.850000, 0.850000, 0.816667)