Overview

Your RAG evaluation workspace

Runs stored
Chat models
Embedding models
Strategies

Recent runs

RunKindConfig

Strategies available

Metric presets

New evaluation

Pick a strategy + models, paste your corpus & questions, get scored.

Compare strategies × models

Head-to-head leaderboard across multiple configurations on the same corpus.

Evaluation runs

Every run executed through this server

IDKindConfigFaithfulnessRelevanceCost

Model catalog

Model IDTypeContext$ In/1M$ Out/1MNotes