Track contract
{{ track.name }}
{{ track.algorithm.summary }}
Fixed by maintainers
Agents per evaluation, scoring, seed policy, and execution defaults.
Chosen for each evaluation
Benchmark, agent, GPU type, algorithm, parallel evaluations, and execution placement.
Default evaluation algorithm
{{ track.algorithm.label }}
{{ track.algorithm.key }}
Evaluation replica
1
User-scaled independently
{% if track.algorithm.shared_compute %}
Reusable compute
1
Shared across task batches
Active agents{{ track.agents_per_evaluation }}Per evaluation
Prepared env pool{{ track.environment_pool_size }}{{ track.algorithm.environment_pool_factor }}× agent slots
Isolated task slots
{{ track.agents_per_evaluation }}
Fresh compute + VM per task
{% endif %}
Resources
- Default GPU
- {{ track.gpu or "None" }}
- Agents / evaluation
- {{ track.agents_per_evaluation }}
- Prepared env pool
- {{ track.environment_pool_size }}
Scoring
- Success bar
- {{ "%.0f" | format(track.success_percent) }}%
- Ranking
- {{ track.ranking_rule }}
- Failed task charge
- {{ track.failure_charge }}
Sampling
- Runs / task
- {{ track.runs_per_task }}
- Seed policy
- {{ track.seed_policy }}
- Leaderboard role
- {{ "Reference only" if track.reference_only else "Competitive" }}
Compatible execution
{% for topology in track.supported_topologies %}
{% else %}
{{ topology.label }}
{% endfor %}
{{ topology.key }}No registered execution topology provides this algorithm's required capabilities.
{% endif %}Versioned Python module
{{ track.algorithm.source_file }}
{% if track.algorithm.aliases %}Aliases: {{ track.algorithm.aliases | join(", ") }}
{% endif %}The algorithm code is included in the run's measurement hash.
Track server configurationJSON
{{ track.server_config | tojson(indent=2) }}
Algorithm source{{ track.algorithm.source_file }}
{{ track.algorithm.source }}