Metadata-Version: 2.4
Name: raggate
Version: 0.1.3
Summary: A thin, CI-gated evaluation gate for RAG & LLM systems: golden set, LLM-judge (or heuristic) scorers, and band-based quality gates.
Project-URL: Homepage, https://github.com/abhay23-AI/raggate
Project-URL: Source, https://github.com/abhay23-AI/raggate
Project-URL: Issues, https://github.com/abhay23-AI/raggate/issues
Project-URL: Changelog, https://github.com/abhay23-AI/raggate/blob/main/CHANGELOG.md
Project-URL: LinkedIn, https://www.linkedin.com/
Author: Abhay Trivedi
License: MIT License
        
        Copyright (c) 2026 Abhay Trivedi
        
        Permission is hereby granted, free of charge, to any person obtaining a copy
        of this software and associated documentation files (the "Software"), to deal
        in the Software without restriction, including without limitation the rights
        to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
        copies of the Software, and to permit persons to whom the Software is
        furnished to do so, subject to the following conditions:
        
        The above copyright notice and this permission notice shall be included in all
        copies or substantial portions of the Software.
        
        THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
        IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
        FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
        AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
        LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
        OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
        SOFTWARE.
License-File: LICENSE
Keywords: ci,evals,evaluation,faithfulness,hallucination,llm,rag,retrieval
Classifier: Development Status :: 3 - Alpha
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Typing :: Typed
Requires-Python: >=3.10
Requires-Dist: pyyaml>=6.0
Provides-Extra: dev
Requires-Dist: pytest>=8.0; extra == 'dev'
Requires-Dist: ruff<0.16,>=0.15; extra == 'dev'
Provides-Extra: openai
Requires-Dist: openai>=1.30; extra == 'openai'
Description-Content-Type: text/markdown

# raggate

A thin, CI-gated evaluation gate for RAG & LLM systems. Point it at a golden set, pick pass/warn/fail bands, and it fails your build when answer quality regresses. Runs with **no API key** (lexical heuristic scorers) and upgrades to **LLM-as-judge** scoring when you set `OPENAI_API_KEY`.

[![PyPI](https://img.shields.io/pypi/v/raggate.svg)](https://pypi.org/project/raggate/)
[![eval-gate](https://github.com/abhay23-AI/raggate/actions/workflows/eval.yml/badge.svg)](https://github.com/abhay23-AI/raggate/actions/workflows/eval.yml)
[![python](https://img.shields.io/pypi/pyversions/raggate.svg)](https://pypi.org/project/raggate/)
[![license: MIT](https://img.shields.io/pypi/l/raggate.svg)](https://github.com/abhay23-AI/raggate/blob/main/LICENSE)

**Fail your CI when RAG/LLM answer quality drops below a threshold.** The part RAGAS and DeepEval leave out: a hard **pass / warn / fail gate** you drop straight into a pull request.

It is deliberately small. It does not replace RAGAS or DeepEval — it is the drop-in gate you wire into GitHub Actions in five minutes, and it composes with them (see [How it compares](#how-it-compares)).

## Why

Most teams test their prompts by eyeballing a few answers and shipping. Then real users arrive and quality quietly regresses with no alarm. `raggate` treats RAG quality like any other production concern: measured against a versioned golden set, and gated in CI.

## Quickstart

```bash
pip install raggate            # core (heuristic mode, no key needed)
# or, for LLM-as-judge scoring:
pip install "raggate[openai]"

raggate init          # scaffolds ./evals (golden set, gates.yaml, target.py)
raggate run           # scores the golden set, prints a banded report
raggate gate          # same, but exits non-zero when a KPI metric FAILs
```

<!-- Demo GIF: render docs/demo.tape with `vhs docs/demo.tape` (see docs/demo.tape),
     commit docs/demo.gif, then add:  ![raggate in CI](docs/demo.gif)  right here.
     A 15s GIF above the fold is the single biggest README click-through win. -->

`raggate run` scores a tiny built-in example immediately:

```
  raggate — 6 case(s) · judge: heuristic
  ──────────────────────────────────────────────────────────────
  METRIC                  SCORE   TARGET   BAND
  ──────────────────────────────────────────────────────────────
  faithfulness            1.000    0.900   PASS
  answer_relevancy        0.715    0.800   WARN
  citation_support        1.000    0.850   PASS
  context_coverage        1.000    0.800   PASS
  answer_correctness      0.627    0.850   WARN
  ──────────────────────────────────────────────────────────────
  heuristic mode — lexical proxies, informational only (never blocks).
```

## Point it at your system

`raggate` calls one function you own. Edit `evals/target.py`:

```python
from my_app.rag import answer   # your existing RAG entrypoint

def run(question: str) -> dict:
    result = answer(question)
    return {
        "answer": result.text,
        "contexts": result.retrieved_chunks,   # list[str] — for coverage & faithfulness
        "citations": result.citations,         # optional
    }
```

## The metrics (and what they actually measure)

Two metric names are chosen to match **what this kit computes**, which is a cheaper proxy than the same-sounding RAGAS metric — so the name does not overpromise. The LLM-judge path is a pointwise rubric score; the heuristic path is a lexical proxy and is **informational only** (never gates).

| Metric | Question | LLM-judge path | Heuristic (no key) | Standard it approximates |
|---|---|---|---|---|
| `faithfulness` | Are the answer's claims grounded in the context? (hallucination) | rubric 0–1 | fraction of sentences covered by context | [RAGAS faithfulness](https://docs.ragas.io/en/stable/concepts/metrics/available_metrics/faithfulness/) (claim-level NLI) |
| `answer_relevancy` | Does the answer address the question? | rubric 0–1 | question-token coverage (weak — can be fooled by echoing) | [RAGAS response relevancy](https://docs.ragas.io/en/stable/concepts/metrics/available_metrics/answer_relevance/) (gen-question + cosine) |
| `citation_support` | Do citations overlap the cited context? | — (overlap only) | citation↔context token coverage | [ALCE citation precision/recall](https://arxiv.org/abs/2305.14627) (NLI) — this is a weaker overlap proxy |
| `context_coverage` | Did retrieval surface the needed tokens? | — (deterministic) | reference-token coverage by retrieved context | [RAGAS context recall](https://docs.ragas.io/en/stable/concepts/metrics/available_metrics/context_recall/) (claim attribution) — **not the same**; renamed to be honest |
| `answer_correctness` | Does the answer match ground truth? | rubric 0–1 | token-F1 vs expected | [RAGAS answer correctness](https://docs.ragas.io/en/stable/concepts/metrics/available_metrics/answer_correctness/) (claim-F1 + cosine) |

Verify a metric name against its source before quoting a number — `context_coverage` and `citation_support` are intentionally *not* the RAGAS metrics.

## Bands, not a single threshold

LLM-as-judge scores are non-deterministic (sampling, position bias, judge-model dependence). A lone `>= 0.80` gate flaps on that noise and teams learn to ignore it. Bands encode intent — edit `evals/gates.yaml`:

```yaml
judge:              # pin the judge — scores are only comparable against the same one
  model: gpt-4.1-mini
  runs: 1           # >1 averages several calls per rating to damp variance
  temperature: 0.0

metrics:
  faithfulness:     { kpi: true, fail: 0.75, warn: 0.85, target: 0.90 }
  context_coverage: { kpi: true, fail: 0.60, warn: 0.70, target: 0.80 }
```

`FAIL` blocks the merge (KPI metrics only). `WARN` reports but passes. Diagnostic (non-KPI) metrics inform without gating. Gating is only enforced in **openai mode** (a judge is configured); **heuristic mode never blocks**. Bands *absorb* variance — they don't remove it — so pin the judge and raise `runs` for high-variance metrics.

**Defaults are starting points for general-purpose RAG, not domain-calibrated.** `faithfulness` carries the highest bar because it is the hallucination guardrail. For regulated domains, use the shipped high-stakes profile — Stanford RegLab found commercial legal RAG tools hallucinate [17–33%](https://reglab.stanford.edu/publications/hallucination-free-assessing-the-reliability-of-leading-ai-legal-research-tools/) of the time, so general-purpose faithfulness/citation floors are unsafe there (note: raggate's `faithfulness` measures grounding in retrieved context, not legal correctness — a distinct concern, but the direction holds):

```bash
raggate gate --gates evals/gates.high-stakes.yaml   # faithfulness/citation floors ≥ 0.90
```

## GitHub Action

raggate ships as a composite Action — the whole gate in three lines:

```yaml
# .github/workflows/rag-quality.yml
name: rag-quality
on: [pull_request]
jobs:
  gate:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: abhay23-AI/raggate@v1
        with:
          openai-api-key: ${{ secrets.OPENAI_API_KEY }}   # omit for heuristic mode
```

The score table is written to the **GitHub Actions run summary** automatically (via `$GITHUB_STEP_SUMMARY`), so you see pass/warn/fail on the PR checks page without opening logs. Inputs: `dir` (default `evals`), `gates`, `openai-api-key`, `python-version`, `version`.

Prefer to run it yourself? It's just the CLI:

```yaml
      - run: pip install "raggate[openai]"
      - run: raggate gate
        env: { OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }} }
```

Without the `OPENAI_API_KEY` secret the gate runs in heuristic mode (informational, never blocks). Add the secret to make it a hard quality gate.

## How it compares

raggate is a *gate*, not a framework. Use it with the others, not instead of them.

| | **raggate** | RAGAS | DeepEval | promptfoo |
|---|:--:|:--:|:--:|:--:|
| Primary job | **CI pass/fail gate** | metric library | metric library + asserts | prompt eval + red-team |
| Runs with no API key | ✅ (heuristic) | ⚠️ limited | ⚠️ limited | ✅ |
| pass / **warn** / fail bands | ✅ | ✗ | threshold | threshold |
| One-line GitHub Action | ✅ | ✗ | ✗ | ✅ |
| Breadth of metrics | 🔸 focused (5) | ✅ many | ✅ **most** | ✅ many |
| Dashboards / tracing | ✗ | ✗ | ✅ (SaaS) | 🔸 |
| Dependency weight | **tiny** (1 dep) | medium | medium | medium |

Where they're stronger: **RAGAS/DeepEval** have far more metrics and DeepEval has a hosted dashboard; **promptfoo** does prompt comparison and red-teaming. If you want depth or dashboards, use those for scoring — raggate is the thin gate that turns any of them into a merge-blocking check.

## Prior art

`raggate` is inspired by these; if you want deep metrics or dashboards, use them — raggate is the thin gate that composes with them.

- **[RAGAS](https://github.com/explodinggradients/ragas)** (Apache-2.0) — the reference library for RAG-specific metrics. raggate's LLM metrics approximate its constructs with pointwise judges; its faithfulness/context-recall definitions are the standard.
- **[DeepEval](https://github.com/confident-ai/deepeval)** (Apache-2.0) — the broadest metric set and pytest-native gating. If you want assertions inside your test suite, use it.
- **[promptfoo](https://github.com/promptfoo/promptfoo)** (MIT) — declarative eval + CI action with RAG assertions.
- **[TruLens](https://github.com/truera/trulens)**, **[Arize Phoenix](https://github.com/Arize-ai/phoenix)** (Elastic-2.0, source-available), **[MLflow LLM eval](https://mlflow.org/docs/latest/genai/eval-monitor/)** — tracing/observability platforms.

## Design notes

1. Retrieval quality (`context_coverage`) is measured separately from generation — most "the LLM is dumb" bugs are retrieval misses.
2. The golden set is versioned and small. Ten sharp cases you trust beat a thousand you don't; grow it every time a bug reaches production.
3. Gating is enforced only in openai mode; heuristic mode (no key) is a smoke test and never blocks. Note `citation_support` and `context_coverage` are deterministic overlap metrics with no LLM path — they still gate in openai mode.

## Roadmap

- [ ] Reranker-lift metric (retrieval score before/after rerank)
- [ ] LLM/NLI path for `citation_support` (toward ALCE-style precision/recall)
- [ ] Adversarial / prompt-injection test pack
- [ ] Multi-run majority (not just mean) aggregation
- [ ] HTML report + PR comment bot

## Contributing

Issues and PRs welcome — see [CONTRIBUTING.md](https://github.com/abhay23-AI/raggate/blob/main/CONTRIBUTING.md). Every metric must run in heuristic mode (no API key) and ship with a test.

## Maintainer

Built and maintained by **Abhay Trivedi** — I build production-grade AI systems and help teams put real eval gates around production RAG/LLM pipelines. Open an issue for the tool, or reach me on [LinkedIn](https://www.linkedin.com/) if you want help wiring evals into your stack.

## License

MIT © Abhay Trivedi. See [LICENSE](https://github.com/abhay23-AI/raggate/blob/main/LICENSE).
