03 / benchmark results

质量与执行结果

{{ run.facets.bundle_label }}

各类任务表现

Q01 · 通过 / 有效执行
{% for category in run.dashboard.categories %}
{{ category.label }}{{ category.rate.value|percent }}通过 {{ category.rate.numerator }} / 有效执行 {{ category.rate.denominator }}{% if category.rate.value is none %} · {{ category.rate.reason or '质量未完整观测' }}{% elif category.rate.excluded or category.unscorable %} · 排除 {{ category.rate.excluded }} · 无法评分 {{ category.unscorable }}{% endif %}
{% else %}

本轮没有适用的分类质量记录。

{% endfor %}

每类保留原分母。质量率不是完成率;执行完成后仍可能答错。六类分别报告,不合成总分。

逐题结果分布

{{ run.requests|length }} 个计划项

按冻结题序排列。选择一个格子查看对应答案。

{% for row in run.requests %}{% endfor %}
质量通过答案未通过未评分 / 未执行
{% for state, label in [('completed','完成'),('failed','模型失败'),('cancelled','取消'),('invalid','无效'),('not_executed','未执行')] %}
{{ run.summary.counts[state] }}{{ label }}
{% endfor %}

答案未通过 {{ run.dashboard.answer_failures }} 项。五种执行终态分别保存,筛选不会更改成绩。

04 / response timing

响应耗时

客户端时间 · 单次观测

逐题请求耗时

{{ run.requests|length }} 个计划项 · 秒
{{ run.dashboard.latency.p50|milliseconds }} s完成请求 P50
{{ run.dashboard.latency.p95|milliseconds }} s完成请求 P95{% if run.dashboard.latency.p95_exploratory %} · 探索性{% endif %}
{{ run.dashboard.first_answer.p50|milliseconds }} s首次答案到达 P50 · {{ run.dashboard.first_answer.sample_count }} 个样本
{{ run.dashboard.latency.sample_count }}有效耗时样本 · 排除 {{ run.dashboard.latency.excluded }}
{% for row in run.requests %}{% if row.latency_ms is not none %}第 {{ row.ordinal }} 题 · {{ row.case_id }} · {{ row.latency_ms|milliseconds }} 秒 · {{ row.state }} / {{ row.quality or '未评分' }}{% endif %}{% endfor %}
题目 1纵轴 0~{{ run.latency_max|milliseconds }} s题目 {{ run.requests|length }}

柱图包含有终止时间的失败请求,棕色表示答案未通过。P50 / P95 仅使用完成请求,按最近秩计算;不足 20 个样本不显示 P95。单次耗时不构成模型间的性能比较资格。