Evaluation report — expense-triage

Customer
northwind
Suite
expense-triage
This run
staging — 2026-08-27T09:45:39.141911+00:00
Reference
staging — 2026-08-27T09:45:38.769454+00:00
Code version
b5870a655c776d75f8be518796bad8acf02fb08a-dirty — uncommitted changes were present, so this run cannot be reproduced from the repository

Did it get worse? Yes

1 check got worse compared with the reference. Every case could be judged. No case is suspended. The configuration is the same as the reference.

Overall

MeasureResultWhat happenedWhy
precision0.700000 / 0.60000020 counted · 0 suspended · 0 not judgedprecision 0.700000 = 7/10 (20 counted, 0 suspended, 0 could not be judged)
accuracy0.800000 / 0.70000020 counted · 0 suspended · 0 not judgedaccuracy 0.800000 = 16/20 (20 counted, 0 suspended, 0 could not be judged)
What got worse (1)
CaseCheckWhat happenedWhy
hotel_overagrees_with_markWent from passing to failing (0.800000 → 0.400000).mean of 5 samples (0.000000, 0.000000, 1.000000, 0.000000, 1.000000)
What could not be judged (0)

Nothing in this section.

What is set aside (0)

Nothing in this section.

What was added or removed (0)

Nothing in this section.

What got better (0)

Nothing in this section.

What stayed the same (21)
CaseCheckWhat happenedWhy
amount_over_capagrees_with_markScore unchanged at 1.000000.mean of 5 samples (1.000000, 1.000000, 1.000000, 1.000000, 1.000000)
book_technicalagrees_with_markScore unchanged at 1.000000.mean of 5 samples (1.000000, 1.000000, 1.000000, 1.000000, 1.000000)
coffee_twoagrees_with_markScore unchanged at 1.000000.mean of 5 samples (1.000000, 1.000000, 1.000000, 1.000000, 1.000000)
courier_urgentagrees_with_markScore unchanged at 1.000000.mean of 5 samples (1.000000, 1.000000, 1.000000, 1.000000, 1.000000)
dinner_clientagrees_with_markScore unchanged at 0.400000.mean of 5 samples (0.000000, 0.000000, 0.000000, 1.000000, 1.000000)
dinner_soloagrees_with_markScore unchanged at 1.000000.mean of 5 samples (1.000000, 1.000000, 1.000000, 1.000000, 1.000000)
duplicate_claimagrees_with_markScore unchanged at 1.000000.mean of 5 samples (1.000000, 1.000000, 1.000000, 1.000000, 1.000000)
hotel_cappedagrees_with_markScore unchanged at 1.000000.mean of 5 samples (1.000000, 1.000000, 1.000000, 1.000000, 1.000000)
late_taxiagrees_with_markScore unchanged at 1.000000.mean of 5 samples (1.000000, 1.000000, 1.000000, 1.000000, 1.000000)
lunch_teamagrees_with_markScore unchanged at 0.600000.mean of 5 samples (1.000000, 0.000000, 1.000000, 0.000000, 1.000000)
meal_no_noteagrees_with_markScore unchanged at 1.000000.mean of 5 samples (1.000000, 1.000000, 1.000000, 1.000000, 1.000000)
monitor_homeagrees_with_markScore unchanged at 1.000000.mean of 5 samples (1.000000, 1.000000, 1.000000, 1.000000, 1.000000)
no_receiptagrees_with_markScore unchanged at 1.000000.mean of 5 samples (1.000000, 1.000000, 1.000000, 1.000000, 1.000000)
parking_airportagrees_with_markScore unchanged at 0.000000.mean of 5 samples (0.000000, 0.000000, 0.000000, 0.000000, 0.000000)
software_seatagrees_with_markScore unchanged at 0.000000.mean of 5 samples (0.000000, 0.000000, 0.000000, 0.000000, 0.000000)
taxi_longagrees_with_markScore unchanged at 0.600000.mean of 5 samples (0.000000, 1.000000, 0.000000, 1.000000, 1.000000)
taxi_shortagrees_with_markScore unchanged at 1.000000.mean of 5 samples (1.000000, 1.000000, 1.000000, 1.000000, 1.000000)
train_standardagrees_with_markScore unchanged at 1.000000.mean of 5 samples (1.000000, 1.000000, 1.000000, 1.000000, 1.000000)
weekend_baragrees_with_markScore unchanged at 1.000000.mean of 5 samples (1.000000, 1.000000, 1.000000, 1.000000, 1.000000)
precisionScore unchanged at 0.700000.precision 0.700000 = 7/10 (20 counted, 0 suspended, 0 could not be judged)
accuracyScore unchanged at 0.800000.accuracy 0.800000 = 16/20 (20 counted, 0 suspended, 0 could not be judged)