Metadata-Version: 2.4
Name: june-bench
Version: 0.0.29
Summary: Reproducible benchmark suite for memory/QA systems — June + pluggable competitors.
Author: Junemind
License: MIT
Project-URL: Homepage, https://github.com/Junemind
Keywords: benchmark,rag,memory,qa,retrieval,evaluation,llm,june
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Operating System :: OS Independent
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: License :: OSI Approved :: MIT License
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Provides-Extra: june-api
Requires-Dist: httpx>=0.24; extra == "june-api"
Provides-Extra: june-local
Requires-Dist: junemind>=0.1.0; extra == "june-local"
Provides-Extra: cognee
Requires-Dist: cognee[evals]; extra == "cognee"
Requires-Dist: fastembed>=0.2; extra == "cognee"
Provides-Extra: stream
Requires-Dist: ijson>=3.2; extra == "stream"
Provides-Extra: all
Requires-Dist: httpx>=0.24; extra == "all"
Requires-Dist: ijson>=3.2; extra == "all"
Provides-Extra: dev
Requires-Dist: pytest>=8; extra == "dev"
Requires-Dist: ijson>=3.2; extra == "dev"
Dynamic: license-file

# june-bench

A pip-installable, **reproducible** benchmark suite for memory / QA systems — **June + pluggable
competitors** — over LoCoMo, LongMemEval, HotpotQA/2Wiki/MuSiQue, and FinanceBench, with the same
data and the same scorer.

```bash
pip install june-bench
june-bench list
june-bench run --system echo --dataset smoke --split smoke    # offline, no key, no download
```

## Reproduce the June vs Cognee head-to-head

One command runs **both** systems over the same HotpotQA open-pool, the same answer model, and the same
judge, and prints a side-by-side with the metered API cost:

```bash
pip install "june-bench[cognee,june-api]"     # bundles cognee + fastembed
june-bench reproduce-h2h --key <YOUR_ACCESS_KEY> --questions 100
```

* **Access key** — June's endpoint is hardware-limited (not yet funded), so runs are key-gated. Request
  one at **access@januraine.ai**; the reply includes your key and this exact command.
* **Same-embedder by default** — Cognee automatically embeds with `bge-large-en-v1.5`, the commodity open
  model June's dense lane uses, so it's a **same-embedder** matched run out of the box (nothing to export).
  This embedder is a disclosed benchmark parameter, not June's moat; pass `--embedder <id>` to swap it.
* **You bring an OpenRouter key** (prompted) — it pays for *both* systems' gpt-4o answers (~$21 for the
  chain-of-thought tier at n=100); the host never holds or pays for it.
* Cognee runs **locally** (needs RAM + a one-time ~1.3 GB fastembed download); June answers over its
  endpoint. The command batches the pool upload, blocks the $90 Opus-on-Cognee path, and meters real cost.

`june-bench reproduce` runs the June-only HotpotQA number the same way; `reproduce-retrieval` scores
June's recall@k/nDCG/MRR. All three are plain-language and need no `JUNE_BENCH_*` env vars.

A benchmark is `run(system, dataset) → records → score`. Two typed ports are the only extension
points:

* **`System`** — the thing benchmarked. `JuneApiSystem` (default; a thin HTTP client to June's
  `/v1/answer`, so **no June source is shipped**), `JuneLocalSystem` (`[june-local]` extra; a
  source-protected compiled wheel), `CogneeSystem` (`[cognee]` extra), or any future system as one
  adapter.
* **`Dataset`** — what it runs on. The four benchmarks behind a registry.

The scorer is the canonical SQuAD/HotpotQA EM/F1 + selective-accuracy/coverage/cost — Cognee-comparable.
Tiny **smoke fixtures ship in the wheel** (offline wiring proof); full splits are **fetched, sha-verified,
from a pinned release**. No score is ever baked into the package — every result row records
dataset + scorer + system + model + cost, so a published number is reproducible by a stranger.

Every result row records dataset + scorer + system + model + cost, and no score is baked into the
package — so a published number is reproducible by a stranger, with the exact command above.
