Jev bench results (SAMPLE: dry-run stub data)
SAMPLE: these numbers come from the offline dry run's stub agent and loopback provider. They are not a bench result.
Stop: all 2 pairs recorded
Before: no JevAfter: with Jev
Accuracy: share of items correctSpeed: items per minute (items / summed wall time)Time consumed: total wall timeHeadline (complete pairs)
| pairs | 2 |
|---|
| accuracy_a | 1 |
|---|
| accuracy_b | 1 |
|---|
| b_only_correct | 0 |
|---|
| a_only_correct | 0 |
|---|
| mcnemar_p | 1 |
|---|
| wall_s_median_difference | 0.3797 |
|---|
| sign_test_p | 0.5 |
|---|
| b_runs_without_jev_answer | 0 |
|---|
Runs
| run | tool | status | use gate | answer | correct | wall s | Jev calls |
|---|
| dry-screen-01.A | jev_screen | ok | n/a (arm has no Jev) | clean | True | 0.02 | 0 |
| dry-screen-01.B | jev_screen | ok | used Jev | clean | True | 0.38 | 1 |
| dry-verify-01.A | jev_verify | ok | n/a (arm has no Jev) | verified | True | 0.02 | 0 |
| dry-verify-01.B | jev_verify | ok | used Jev | verified | True | 0.42 | 1 |