Metadata-Version: 2.4
Name: turkish-rag-eval
Version: 0.2.0
Summary: Measure which parts of a Turkish RAG pipeline pay off: chunking, stemming, embedding model, fusion.
Author: Rızgar Ozan
License-Expression: MIT
Project-URL: Homepage, https://github.com/RizgarOzan/turkish-rag-eval
Project-URL: Dataset, https://huggingface.co/datasets/RizgarOzan/turkish-rag-eval
Project-URL: Issues, https://github.com/RizgarOzan/turkish-rag-eval/issues
Keywords: turkish,retrieval,rag,bm25,benchmark,evaluation
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Science/Research
Classifier: Natural Language :: Turkish
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Text Processing :: Indexing
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
License-File: NOTICE.md
Requires-Dist: numpy>=1.24
Requires-Dist: rank-bm25>=0.2.2
Requires-Dist: requests>=2.31
Provides-Extra: dense
Requires-Dist: sentence-transformers>=3.0; extra == "dense"
Requires-Dist: torch>=2.0; extra == "dense"
Provides-Extra: charts
Requires-Dist: matplotlib>=3.7; extra == "charts"
Provides-Extra: llm
Requires-Dist: anthropic>=0.40; extra == "llm"
Provides-Extra: dev
Requires-Dist: pytest>=7.4; extra == "dev"
Provides-Extra: all
Requires-Dist: turkish-rag-eval[dense]; extra == "all"
Requires-Dist: turkish-rag-eval[charts]; extra == "all"
Requires-Dist: turkish-rag-eval[llm]; extra == "all"
Requires-Dist: turkish-rag-eval[dev]; extra == "all"
Dynamic: license-file

# turkish-rag-eval

[![validate](https://github.com/RizgarOzan/turkish-rag-eval/actions/workflows/validate.yml/badge.svg)](https://github.com/RizgarOzan/turkish-rag-eval/actions/workflows/validate.yml)
[![licence: MIT + CC BY-SA 4.0](https://img.shields.io/badge/licence-MIT%20%2B%20CC%20BY--SA%204.0-blue)](https://github.com/RizgarOzan/turkish-rag-eval/blob/main/NOTICE.md)
[![dataset on Hugging Face](https://img.shields.io/badge/%F0%9F%A4%97%20dataset-RizgarOzan%2Fturkish--rag--eval-yellow)](https://huggingface.co/datasets/RizgarOzan/turkish-rag-eval)

**Which parts of a RAG pipeline actually earn their cost on an agglutinative
language?** Three chunking strategies × four retrievers × six embedding
models, measured on a hand-labelled Turkish gold set, with confidence
intervals and a cost column — then pointed at your own documents.

```bash
pip install git+https://github.com/RizgarOzan/turkish-rag-eval   # not on PyPI yet
turkish-rag-eval run --corpus ./belgelerim --gold ./sorular.json
turkish-rag-eval report
```

## Four findings

**1. The embedding model matters more than any pipeline choice.** Every model
trained for retrieval beats stemmed BM25 (0.494) on its own. The best,
[Mursit-Large-TR-Retrieval](https://huggingface.co/newmindai/Mursit-Large-TR-Retrieval),
reaches **0.781** nDCG@10 on hierarchical chunks. A small E5 the same size as
the default goes from 0.501 to 0.642 with nothing else changed, at about the
same speed (22 ms against 23 ms). The popular default model is the weak link, not dense retrieval.
See the [Leaderboard](#leaderboard).

**2. Turkish stemming is the cheapest real win.** Truncating tokens to a
5-character prefix before BM25 lifts nDCG@10 by 23–29% on every chunking
strategy, and a paired bootstrap puts every one of those gains clear of zero
(`hierarchical +0.111, 95% CI [+0.026, +0.204]`). "diyabet", "diyabetin",
"diyabete", "diyabetli" are four surface forms of one concept; an unstemmed
index almost never matches the query's form.

**3. The top of the table is a tie, and the tie decides on cost.** With 58
queries, a paired bootstrap cannot separate the best configuration from the
other two hybrid ones; the remaining nine are measurably worse. So the
real choice at the top is price: `sentence + hybrid_rrf` gives the same
quality at 27 ms instead of 37 ms.

<picture>
  <source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/RizgarOzan/turkish-rag-eval/main/docs/charts/ndcg-intervals-dark.png">
  <img alt="nDCG@10 per configuration with 95% bootstrap intervals; the best cannot be told apart from the other two hybrid configurations, the remaining nine are measurably worse" src="https://raw.githubusercontent.com/RizgarOzan/turkish-rag-eval/main/docs/charts/ndcg-intervals.png">
</picture>

**4. The intuitive confidence signal is the useless one.** For deciding when
*not* to answer, the obvious measure — how far ahead the top hit is — is worse
than answering everything (0.25 selective accuracy against a 0.47 baseline).
RRF fuses ranks as `1/(60+rank)`, so the top-two gap is ~2% on every query,
confident or not. The dense retriever's raw cosine works; the margin does not.

<picture>
  <source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/RizgarOzan/turkish-rag-eval/main/docs/charts/abstention-dark.png">
  <img alt="Coverage against selective accuracy for two confidence signals; the top-1 margin performs worse than answering everything" src="https://raw.githubusercontent.com/RizgarOzan/turkish-rag-eval/main/docs/charts/abstention.png">
</picture>

Longer write-up of the first result:
[English](https://github.com/RizgarOzan/turkish-rag-eval/blob/main/docs/blog/2026-09-19-bm25-turkish-en.md) ·
[Türkçe](https://github.com/RizgarOzan/turkish-rag-eval/blob/main/docs/blog/2026-09-19-bm25-turkish-tr.md).

**Contents:** [Results](#results) · [Leaderboard](#leaderboard) ·
[Why not an existing benchmark?](#why-not-an-existing-benchmark) ·
[Your own corpus](#your-own-corpus) · [Running it](#running-it) ·
[Gold set](#gold-set) · [Status](#status) · [Limits](#limits) · [Contribute](#contribute) ·
[Design notes](https://github.com/RizgarOzan/turkish-rag-eval/blob/main/docs/design.md)

## Results

58 queries, default embedding model
`paraphrase-multilingual-MiniLM-L12-v2`, CPU only. Intervals are 95% bootstrap
over queries; `turkish-rag-eval report` regenerates this table.

| Chunking | Retriever | nDCG@10 | 95% CI | R@5 | MRR | P95 |
|---|---|---|---|---|---|---|
| **hierarchical** | **hybrid_rrf** | **0.613** | [0.508, 0.716] | 0.690 | 0.559 | 37 ms |
| sentence | hybrid_rrf | 0.558 | [0.453, 0.662] | 0.621 | 0.500 | 27 ms |
| fixed | hybrid_rrf | 0.517 | [0.404, 0.629] | 0.578 | 0.484 | 31 ms |
| sentence | bm25_stem5 | 0.510 | [0.403, 0.621] | 0.638 | 0.462 | 4 ms |
| hierarchical | dense | 0.501 | [0.396, 0.606] | 0.569 | 0.441 | 30 ms |
| hierarchical | bm25_stem5 | 0.494 | [0.386, 0.605] | 0.569 | 0.444 | 8 ms |
| fixed | bm25_stem5 | 0.476 | [0.371, 0.585] | 0.526 | 0.421 | 4 ms |
| sentence | dense | 0.461 | [0.356, 0.567] | 0.552 | 0.410 | 21 ms |
| fixed | dense | 0.446 | [0.341, 0.555] | 0.491 | 0.402 | 25 ms |
| sentence | bm25_nostem | 0.411 | [0.308, 0.515] | 0.483 | 0.356 | 4 ms |
| fixed | bm25_nostem | 0.387 | [0.290, 0.487] | 0.414 | 0.319 | 3 ms |
| hierarchical | bm25_nostem | 0.383 | [0.285, 0.482] | 0.500 | 0.319 | 6 ms |

`turkish-rag-eval report` does not stop at the table — it names the cheapest
configuration the data cannot separate from the best:

```
Recommended: sentence + hybrid_rrf
  0.558 ndcg@10 against 0.613 for hierarchical + hybrid_rrf, a gap of 0.056
  that a paired bootstrap over 58 shared queries cannot distinguish from zero.
  It answers in 27 ms at P95 against 37 ms.
```

Comparisons are **paired**: both systems answered the same queries, so
resampling per-query differences removes the variance of some queries being
harder than others. It matters, and this data shows it in the sharpest way.
`hierarchical + hybrid_rrf` beats `fixed + hybrid_rrf` by 0.097 and beats
`sentence + bm25_stem5` by 0.104 — yet only the **larger** gap clears zero,
because its per-query differences are so much steadier (sd 0.34 against 0.40).
Ranking by the gap alone gets this backwards.

Hierarchical chunking only helps the dense retriever, and fixed-size chunking
— the most common default — lost on every retriever. More in the
[design notes](https://github.com/RizgarOzan/turkish-rag-eval/blob/main/docs/design.md#what-else-the-numbers-say).

### Generation: groundedness

The R above feeds a G: the 58 questions answered from the top 5
`hierarchical + bm25_stem5` passages by `groq:openai/gpt-oss-120b`, judged by
a model from another family, `nvidia:nvidia/nemotron-3-super-120b-a12b`, at
temperature 0 on `2026-09-30`. The gold span was among the passages for
33 of 58 questions, and the two groups are scored apart:

| Span retrieved (33) | |
|---|---|
| answered when the span was retrieved | 93.9% |
| answer rests on the passages | 96.8% |
| answer conveys the gold span | 96.8% |

| Span not retrieved (25) | |
|---|---|
| declined when the span was not retrieved | 80.0% |
| answered anyway | 20.0% |
| of those, answer is right | 60.0% |
| hallucinated (not in the passages) | 4.0% |

The generator mostly declines when the evidence is missing, and when it does
answer, the answer is usually stated in some other retrieved passage — one of
25 is not. The first scoring rule called every answer without the span a
hallucination and reported 20%; reading the five by hand showed four stated in
a retrieved passage, so answers without the span now go to the judge too
([why](https://github.com/RizgarOzan/turkish-rag-eval/blob/main/docs/design.md#groundedness)).
16 verdicts were checked by hand; 14 agreed. Of the other two, the judge was
too strict once (an answer inverted "not recommended unless below 7 g/dl") and
too lenient once (it accepted "hyperuricaemia" as the cause of gout).

This row depends on hosted APIs and cannot be reproduced offline; the models
may change behind the same name. Only the rates are committed
([`results/groundedness.json`](https://github.com/RizgarOzan/turkish-rag-eval/blob/main/results/groundedness.json)) —
the free API terms do not allow redistributing raw model output.

To rerun it with other models, name them `groq:`, `nvidia:` or `openai:`
(a bare name goes to Anthropic). Keys can sit in a `.env` file (gitignored) in
the directory you run from; `openai:` follows `OPENAI_BASE_URL`, so any OpenAI-compatible
endpoint, a local router included, works.

## Leaderboard

Dense `nDCG@10` per chunking strategy, the hybrid (dense + stemmed BM25, RRF)
on hierarchical chunks, and the cost of each model on one 16-thread CPU,
measured 2026-09-18. Stemmed BM25 alone scores 0.476 / 0.510 / 0.494.
`turkish-rag-eval leaderboard` rebuilds this from `results/models/`.

| Model | Params | Turkish-only | Dense fixed | Dense sentence | Dense hierarchical | Hybrid hierarchical | Embed corpus | Query P95 |
|---|---|---|---|---|---|---|---|---|
| paraphrase-multilingual-MiniLM-L12-v2 (default) | 118 M | no | 0.446 | 0.461 | 0.501 | 0.607 | 2 min | 23 ms |
| [emrecan/bert-base-turkish-cased-mean-nli-stsb-tr](https://huggingface.co/emrecan/bert-base-turkish-cased-mean-nli-stsb-tr) | 111 M | yes | 0.408 | 0.431 | 0.497 | 0.654 | 4 min | 43 ms |
| [intfloat/multilingual-e5-small](https://huggingface.co/intfloat/multilingual-e5-small) | 118 M | no | 0.654 | 0.644 | 0.642 | 0.639 | 3 min | 22 ms |
| [intfloat/multilingual-e5-base](https://huggingface.co/intfloat/multilingual-e5-base) | 278 M | no | 0.631 | 0.677 | 0.668 | 0.648 | 10 min | 49 ms |
| [BAAI/bge-m3](https://huggingface.co/BAAI/bge-m3) | 568 M | no | 0.766 | 0.767 | — | — | > 45 min | — |
| **[newmindai/Mursit-Large-TR-Retrieval](https://huggingface.co/newmindai/Mursit-Large-TR-Retrieval)** | 404 M | yes | **0.746** | **0.740** | **0.781** | 0.673 | 35 min | 214 ms |

The two Turkish-only models are the most-downloaded Turkish entries in the
Hugging Face `sentence-similarity` category. `bge-m3` was stopped after 45
minutes, before the hierarchical chunks; its two numbers come from that
partial run. `google/embeddinggemma-300m` is gated behind a licence click and
was not run. Per-query results for every completed model are in
`results/models/`.

> These five rows were measured together on one machine before v0.1.0, which
> is why the timing columns are comparable with each other and not with the
> main table above. Their `dense` columns are unaffected by the stable
> tie-break, but each `Hybrid hierarchical` figure will move by roughly +0.006
> when the model is re-run — the default row's went 0.607 → 0.613.
> `turkish-rag-eval leaderboard --check` reports them as missing provenance
> until then.

**The top of this table is settled; the middle is not.** Dense hierarchical,
with 95% bootstrap intervals over the 58 queries: MiniLM 0.501 [0.396, 0.606],
emrecan 0.497 [0.394, 0.601], e5-small 0.642 [0.542, 0.738], e5-base 0.668
[0.568, 0.765], Mursit 0.781 [0.701, 0.856]. The intervals overlap, but paired
over the same queries Mursit beats the runner-up e5-base by +0.113
[+0.029, +0.202]. The two E5 models cannot be told apart (+0.026
[-0.039, +0.094]), and neither can MiniLM and emrecan (+0.004
[-0.123, +0.135]).

**Hybrid fusion only pays for a weak dense model.** RRF lifts the small
default by +0.106 but pulls Mursit down from 0.781 to 0.673, a paired loss of
-0.108 [-0.193, -0.030]; for the two E5 models it makes no measurable
difference. "Turkish-only" is not enough either: the `emrecan` model was
trained for sentence similarity and truncates input at 75 tokens.

To submit a model, open a pull request adding `results/models/<org>__<name>/`
(`turkish-rag-eval run --model <org>/<name>`, then `leaderboard --check`).
CI re-derives every metric from the per-query relevance arrays committed
beside it, and checks the harness version and corpus fingerprint.

## Why not an existing benchmark?

MTEB-style retrieval benchmarks, TR-MTEB included, score an embedding model on
passages that are already split. They answer "which model?", not "which
chunker, is Turkish stemming worth it, does a hybrid help, and what does each
cost on a CPU?". This harness keeps the articles whole, lets every chunker cut
them its own way, and judges each chunk by the answer span, so pipeline
choices are compared on the same labels. For a model-only comparison the same
data exports to the BEIR layout MTEB reads (`turkish-rag-eval export-hf`),
published as
[RizgarOzan/turkish-rag-eval](https://huggingface.co/datasets/RizgarOzan/turkish-rag-eval)
(the 58 human questions, plus all 296 under the `full-*` configs, each marked
`human` or `llm-draft`); adding it to MTEB is proposed in
[embeddings-benchmark/mteb#5536](https://github.com/embeddings-benchmark/mteb/issues/5536).

## Your own corpus

The harness is not tied to its own articles. Point it at a folder of `.txt` or
`.md` files, a BEIR directory, or a JSON corpus:

```bash
turkish-rag-eval run --corpus ./belgelerim --gold ./sorular.json \
                     --retrievers bm25_stem5 bm25_nostem
```

A BM25-only run needs no model download and no torch — enough to answer "which
chunker, and is stemming worth it on my documents" from the base install.

**No labelled questions?** That is the real wall, and the reason most
benchmarks only ever measure themselves. `bootstrap` drafts a starting point (the name means drafting a gold set here,
not the statistical resampling above):

```bash
export ANTHROPIC_API_KEY=...
turkish-rag-eval bootstrap ./belgelerim --out gold-draft.json --per-doc 3
```

It runs the same procedure the contributed questions in this repository used.
One pass writes questions and marks the answer span; a second pass sees only
the document and the questions and marks the span again. Where the passes
disagree the item is written as `"review": "needs-human"` and **never loads**
until a person settles it. Spans that are not verbatim, and questions copied
out of their own answer, are dropped with a reason.

What you get is a draft to review, not a gold set.

## Running it

Python 3.10+. CPU only — no GPU anywhere in this project.

```bash
pip install git+https://github.com/RizgarOzan/turkish-rag-eval                       # metrics, BM25, gold-set tooling
pip install 'turkish-rag-eval[all] @ git+https://github.com/RizgarOzan/turkish-rag-eval'  # + dense retrieval, charts, LLM commands
```

| Command | What it does |
|---|---|
| `run` | every chunking × retriever combination; writes `results/` |
| `report` | intervals, paired comparisons, and a recommendation |
| `charts` | the two figures above, light and dark |
| `leaderboard` | rebuild the model table; `--check` verifies every entry |
| `abstain` | coverage / selective-accuracy curve for the best configuration ([details](https://github.com/RizgarOzan/turkish-rag-eval/blob/main/docs/design.md#abstention)) |
| `groundedness` | score the generation half against the gold spans ([details](https://github.com/RizgarOzan/turkish-rag-eval/blob/main/docs/design.md#groundedness); [results](#generation-groundedness)) |
| `bootstrap` | draft a gold set for your own corpus |
| `agreement` | inter-annotator agreement over the gold set |
| `fetch-corpus` | download the Wikipedia snapshot; `--verify` checks the lock; `--include-drafts` for [all 300 questions](#scoring-all-300) |
| `export-hf` | the BEIR layout MTEB reads |
| `validate` | every gold file's invariants |

From a checkout:

```bash
git clone https://github.com/RizgarOzan/turkish-rag-eval
cd turkish-rag-eval
pip install -e '.[all]'
turkish-rag-eval fetch-corpus     # rebuilds data/raw/corpus.json
turkish-rag-eval run              # results/summary.json + per-query files
python -m pytest -q
```

`run --model <name>` swaps the embedding model; any sentence-transformers
model works. Results for a non-default model go to
`results/models/<org>__<name>/`, so the main table is never overwritten.

## Gold set

| Files | Questions | Labelled by | In the results above |
|---|---|---|---|
| `data/eval/gold.json` (health) | 58 | one human | yes |
| `data/eval/contrib/llm-draft-*.json` (history, geography, astronomy, biology, computing) | 242 | two independent LLM passes | not yet |

The set reached its planned 300 questions on 2026-09-26. Questions a model
drafted are marked `"source": "llm-draft"`. A second model then picked its own
answer span for each one without seeing the first label
(`second_annotation`). Two spans agree when one contains the other or their
token F1 is at least 0.5; agreed items get `"review": "agreed"`, the rest get
`"needs-human"` and are never loaded. Batches 1 and 2 were 30 of 30 agreed,
batch 3 was 29 of 30: for "what are several ribosomes working on one mRNA
called?" the passes picked two different sentences that both name polysomes,
so that question waits for a person. Batch 4 was 27 of 30: the second pass
named the other claimant to the Hungarian throne, took the sentence beside the
lysozyme result instead of the result itself, and answered "cross compilers"
with the bare term where the first took its definition. Batch 5 was 30 of 30,
and every pair is a containment pair: 9 identical, 15 differing only by a
trailing full stop. One of its answers, Robert W. Holley's 1968 Nobel Prize,
was cut to its last clause on 2026-09-27 and labelled again by the second
pass: the sentence splitter breaks after "W.", so no hierarchical chunk held
the whole sentence ([#17](https://github.com/RizgarOzan/turkish-rag-eval/issues/17)). Batch 6 was 30 of 30 on reworded questions whose answers
often run to two sentences; six first-pass spans were cut to one sentence
before merging, because the two-sentence version fitted inside no chunk and
so could never be retrieved. Batch 7 was 30 of 30 again, with 25 identical
spans; one first-pass span was cut to its clause for the same reason. Like
batch 5 its questions mostly point at a single sentence, so its high
agreement says little about harder questions. Batch 8, the last 32, was
written to be harder: 17 first-pass answers ran to two or more sentences.
Seven of those crossed a sentence- or hierarchical-chunk boundary and were
cut, leaving 12 multi-sentence answers. It was still 32 of 32, 29 identical:
two LLMs reading the same article pick the same sentences even when the
answer is long, which is one more reason these drafts need a human pass.

Drafts stay out of every number above until a re-run says otherwise:
`load_gold()` skips them unless called with `include_drafts=True`. Synthetic
test questions are common practice as long as they are declared and their
agreement is measured. This section is that declaration.

The corpus is 54 Turkish Wikipedia articles, 1.09 M characters; 27 of them
answer at least one question and the other 27 are distractors from the same
domain.

| Batch | Questions | Identical | Mean IoU | Mean token F1 | Cohen's κ, fixed / sentence / hierarchical |
|---|---|---|---|---|---|
| 1 — 2026-09-18 (Malazgirt, Kapadokya, Mars, Mitokondri, Linux) | 30 | 16 | 0.816 | 0.878 | 1.00 / 1.00 / 1.00 |
| 2 — 2026-09-20 (İstanbul'un Fethi, Ağrı Dağı, Jüpiter, Fotosentez, İnternet)¹ | 30 | 4 | 0.427 | 0.549 | 0.94 / 0.99 / 0.99 |
| 3 — 2026-09-21 (Çaldıran Muharebesi, Tuz Gölü, Satürn, Ribozom, Unix) | 30 | 18 | 0.829 | 0.872 | 0.92 / 0.96 / 0.96 |
| 4 — 2026-09-24 (Mohaç Muharebesi, Kızılırmak, Venüs, Enzim, Derleyici) | 30 | 15 | 0.747 | 0.792 | 0.96 / 0.96 / 0.96 |
| 5 — 2026-09-25 (Kösedağ Muharebesi, Uludağ, Neptün, RNA, İşletim sistemi) | 30 | 24 | 0.922 | 0.943 | 1.00 / 1.00 / 1.00 |
| 6 — 2026-09-25 (Preveze Deniz Muharebesi, Erciyes, Uranüs, Hemoglobin, Veritabanı) | 30 | 20 | 0.869 | 0.901 | 0.93 / 0.90 / 0.87 |
| 7 — 2026-09-26 (Ankara Muharebesi, Van Gölü, Merkür, DNA, World Wide Web) | 30 | 25 | 0.944 | 0.960 | 0.93 / 0.93 / 0.97 |
| 8 — 2026-09-26 (Niğbolu Muharebesi, Fırat, Ay, Protein, Yapay zekâ) | 32 | 29 | 0.964 | 0.975 | 0.94 / 0.97 / 1.00 |
| **All drafts** | 242 | 151 | 0.816 | 0.860 | 0.95 / 0.97 / 0.97 |

¹ Re-measured 2026-09-26: the Turkish Wikipedia article *Jüpiter* was rewritten that day to correct errors, and five of its spans no longer matched the live text. Three changed only in wording (a comma, "30,003" → "30" seconds, "dört uydu" → "uydular") and were edited in both labels; two changed in substance — the 40,000 km mantle thickness is gone and the Great Red Spot went from "at least 400 years" to "recorded since 1831" — so those two questions were rewritten and labelled again by both passes. The row was 0.419 / 0.540 / 0.96 / 1.00 / 1.00 before.

The passes disagree about how much of a sentence to take, rather than where
the answer is. Two LLMs tend to pick the same sentence, so read this as a
sanity check rather than human agreement. Full discussion in the
[design notes](https://github.com/RizgarOzan/turkish-rag-eval/blob/main/docs/design.md#agreement-between-the-two-passes).

### Scoring all 300

The drafted questions point at 40 articles outside the health snapshot, so
they get a corpus of their own: the 54 health articles plus those 40, 1.89 M
characters, pinned by `data/corpus-full.lock.json`. The published snapshot and
its lock stay as they are.

```bash
turkish-rag-eval fetch-corpus --include-drafts   # data/raw/corpus-full.json
turkish-rag-eval run --include-drafts            # results/full/<model>/
```

Measured 2026-09-27 (re-run after the batch 5 fix; Mursit added 2026-10-01) on 296 questions (the 4 `needs-human` drafts never load),
hierarchical chunks, nDCG@10:

| Retriever | 58 human questions | 296 questions (238 drafted) |
|---|---|---|
| bm25_stem5 | 0.494 | 0.558 |
| MiniLM (default), dense | 0.501 | 0.405 |
| MiniLM (default), hybrid_rrf | 0.613 | 0.578 |
| multilingual-e5-small, dense | 0.642 | 0.568 |
| multilingual-e5-small, hybrid_rrf | 0.639 | 0.642 |
| multilingual-e5-base, dense | 0.668 | 0.646 |
| multilingual-e5-base, hybrid_rrf | 0.648 | 0.664 |
| Mursit-Large-TR-Retrieval, dense | 0.781 | 0.613 |
| Mursit-Large-TR-Retrieval, hybrid_rrf | 0.673 | 0.661 |

On the wider set stemmed BM25 gets stronger and every dense model weaker, so
fusing the two now helps all four models, where on the 58 health questions it
cost the E5 models and Mursit. Mursit drops the most, from 0.781 to 0.613, and
falls below e5-base, so its lead on the human set does not carry over to the
drafted questions yet. Those questions were written by an LLM from one sentence
each; whether that style suits the smaller models better is open until people
have checked them. These rows are not in the tables above: most of the
questions have not been checked by a person yet. The run also names every
question no chunk can answer; a span split by a chunk boundary is the usual
reason. Here that is 17 questions with fixed-size chunks and none with
sentence or hierarchical chunks (the one hierarchical miss was fixed on
2026-09-27, see batch 5 above). Results for all three chunkers are in
[`results/full/`](https://github.com/RizgarOzan/turkish-rag-eval/tree/main/results/full).

## Status

v0.1.0, tagged but not on PyPI yet — install from GitHub as above.

- **Works:** the retrieval harness, [`groundedness`](#generation-groundedness) on the 58 human questions, `report`, `leaderboard --check`, `bootstrap`,
  `agreement`, the Hugging Face export, scoring [all 300 questions](#scoring-all-300)
  on a pinned corpus, and CI that validates every contributed question against
  live Wikipedia.
- **In progress:** human review of the gold set. All 300 planned questions are
  in (58 human, 242 LLM-drafted); the drafts join the results once people have
  reviewed them.
- **Not yet:** EmbeddingGemma, and `groundedness` on the 242 drafted questions.

## Limits

- **58 queries is a small set.** The intervals above are the honest width of
  that; most of the table's ordering is not resolved by this much data. They
  replace the eyeballed rule of thumb this project used to carry — "treat
  differences under roughly 0.05 nDCG as noise" — which was both too strict
  for paired comparisons and too loose for unpaired ones.
- **One annotator for the human set.** The 58 health questions are
  single-annotated, so they have no agreement figure. The 242 drafted questions
  are double-labelled, but by two LLM passes rather than by two people.
- **The corpus is Wikipedia, and so is much of the training data.** Every
  embedding model ranked here was almost certainly trained on Turkish
  Wikipedia. Absolute scores are therefore optimistic; the *comparison*
  between pipeline choices on the same corpus is what this measures. A
  non-Wikipedia domain is the most valuable thing a contributor could add.
- **Six embedding models, one corpus.** `bge-m3` is a partial run and
  EmbeddingGemma is missing. `abstain` uses only the default model.
- **Encyclopaedic text, not clinical text.** Nothing here transfers to a
  clinical setting without re-measurement. No patient data is used anywhere,
  and nothing here is a medical device.
- **The generation half is one generator, one judge, 58 questions.**
  [Groundedness](#generation-groundedness) was measured once, through hosted
  APIs, on the human set only; the judge was checked by hand on 16 verdicts.
- **The embedding model is pinned by name, not by revision.**

## Contribute

The first two limits shrink with every contributor. Adding questions needs no
ML background — pick a Turkish Wikipedia article, write 5–10 paraphrased
questions, and open a pull request with one JSON file. A validator checks each
file against Wikipedia in CI. See [CONTRIBUTING.md](https://github.com/RizgarOzan/turkish-rag-eval/blob/main/CONTRIBUTING.md) (Türkçe
açıklama dahil) and the
[good first issues](https://github.com/RizgarOzan/turkish-rag-eval/issues?q=is%3Aopen+label%3A%22good+first+issue%22).

Submitting an embedding model is one command and a pull request — see
[Leaderboard](#leaderboard).

## Data

Turkish Wikipedia, CC BY-SA 4.0. See [NOTICE.md](https://github.com/RizgarOzan/turkish-rag-eval/blob/main/NOTICE.md).

## Licence

Code MIT ([LICENSE](https://github.com/RizgarOzan/turkish-rag-eval/blob/main/LICENSE)); data under `data/` CC BY-SA 4.0
([NOTICE.md](https://github.com/RizgarOzan/turkish-rag-eval/blob/main/NOTICE.md)).
