Metadata-Version: 2.5
Name: riskval
Version: 0.1.0
Summary: Regulatory validation battery for credit-risk and market-risk models: PD calibration tests, VaR/ES backtests, Basel IRB capital.
Project-URL: Homepage, https://github.com/datasciengine/riskval
Project-URL: Issues, https://github.com/datasciengine/riskval/issues
Project-URL: Changelog, https://github.com/datasciengine/riskval/blob/main/CHANGELOG.md
Author: Murat Sahin
License-Expression: MIT
License-File: LICENSE
Keywords: backtesting,basel,calibration,credit-risk,irb,model-risk,model-validation,value-at-risk
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Financial and Insurance Industry
Classifier: Intended Audience :: Science/Research
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Topic :: Office/Business :: Financial
Classifier: Topic :: Scientific/Engineering :: Mathematics
Classifier: Typing :: Typed
Requires-Python: >=3.10
Requires-Dist: numpy>=1.24
Requires-Dist: pandas>=2.0
Requires-Dist: scipy>=1.10
Provides-Extra: dev
Requires-Dist: build; extra == 'dev'
Requires-Dist: pytest-cov; extra == 'dev'
Requires-Dist: pytest>=8; extra == 'dev'
Requires-Dist: ruff>=0.6; extra == 'dev'
Requires-Dist: scikit-learn; extra == 'dev'
Requires-Dist: statsmodels; extra == 'dev'
Requires-Dist: twine; extra == 'dev'
Description-Content-Type: text/markdown

# riskval

**A regulatory validation battery for credit- and market-risk models.**

Python has no maintained, tested implementation of the statistical tests that
bank supervisors actually apply to internal models. Model risk management teams
rewrite them in Excel or SAS every year. `riskval` is that battery, in one
package, with every number checked against a published reference.

```bash
pip install riskval
```

## Quickstart

```python
>>> import numpy as np
>>> from riskval import validate
>>>
>>> rng = np.random.default_rng(0)
>>> pd_true = rng.uniform(0.01, 0.30, size=4000)   # the true default rates
>>> y = rng.binomial(1, pd_true)                   # who actually defaulted
>>>
>>> report = validate(y, pd_true, seed=0)
>>> report.rejected
[]

```

Now hand the same battery a model that ranks obligors perfectly but prices them
at half the risk:

```python
>>> report = validate(y, pd_true / 2, seed=0)
>>> report.rejected
['Hosmer-Lemeshow', 'Spiegelhalter Z']

```

The ranking is untouched, so every discrimination measure is bit-for-bit
identical to the good model's:

```python
>>> good = validate(y, pd_true, seed=0)
>>> round(report.metrics["gini"], 6) == round(good.metrics["gini"], 6)
True

```

That gap is the whole point. A model can look excellent on AUC and still fail
every test a supervisor runs. And it is not free: put the misestimated PD
through the Basel IRB formula and read off the capital shortfall.

```python
>>> from riskval import rwa_deviation
>>> out = rwa_deviation(pd_true=0.03, pd_hat=0.01, lgd=0.45, ead=1_000_000)
>>> out["direction"]
'understated'
>>> round(out["rwa_relative_deviation"], 4)
-0.2812

```

A grade priced at 1% that actually defaults at 3% carries 28% too little
capital.

## What is in it

| Module | Question | Contents |
| --- | --- | --- |
| `riskval.calibration` | Is the *level* of the PDs right, and is their *spread*? | Exact binomial test, Jeffreys test, Hosmer–Lemeshow, Spiegelhalter's Z, Cox calibration slope and joint test, calibration curve table, `grade_level_report` |
| `riskval.discrimination` | Is the *ordering* right? | AUC with bootstrap CI, Gini, Kolmogorov–Smirnov, Somers' D |
| `riskval.backtest` | Does the VaR model hold up? | Kupiec POF, Christoffersen independence and conditional coverage, Basel traffic light, capital multiplier add-on |
| `riskval.capital` | What does being wrong cost? | Basel IRB risk weights for four asset classes, RWA, `rwa_deviation` |
| `riskval.report` | All of it at once | `validate`, `ValidationReport` |

Dependencies: `numpy`, `scipy`, `pandas`. Nothing else, and nothing proprietary
— no Bloomberg, no Refinitiv, no vendor data.

## The promise: every number traces to a published reference

Nobody installs a package for a formula they could write in twenty lines. They
install it so they do not have to *defend* those twenty lines. So each test is
checked in the test suite against a value someone else published:

| Test | Reference |
| --- | --- |
| Exact binomial | BCBS (1996) Table 1: cumulative binomial probabilities for 250 days at 1%, all nine rows. Plus the tail in exact rational arithmetic via `fractions.Fraction`. |
| Jeffreys | The beta–binomial identity, and quadrature of the posterior density independent of `scipy.special.betainc` |
| Hosmer–Lemeshow | Closed-form arithmetic done by hand, plus a reproduction of Hosmer and Lemeshow's own simulation establishing the null distribution |
| Spiegelhalter Z | Algebraic equivalence of the compact form and the standardised Brier score, plus a Monte Carlo check that the null really is *N*(0,1) |
| Cox calibration slope | A saturated two-point design, whose intercept and slope have a closed form, matched to 1e-10; `statsmodels` for the coefficients, standard errors and log-likelihood; and recovery of a known distortion (predictions pushed to *c* times their true log-odds must return a slope of 1/*c*) |
| AUC, Gini, Somers' D | Hanley and McNeil (1982) Table 1, whose published area of 0.893 is reproduced as the exact rational 1321/1479; cross-checked against `sklearn` and `scipy.stats.mannwhitneyu`, and against brute-force pair counting |
| Kolmogorov–Smirnov | `scipy.stats.ks_2samp` |
| Kupiec POF | Kupiec (1995) Table 1: the 95% non-rejection regions, all fifteen cells, bound by bound |
| Christoffersen | Longhand re-derivation of both likelihood ratios, plus Monte Carlo null distributions |
| Basel IRB risk weights | BCBS (2006) Annex 3 "Illustrative IRB Risk Weights": all 19 PD rows × 4 asset classes, and the formula rebuilt from the regulatory text |
| Basel traffic light | BCBS (1996) Tables 1 and 2 |

If a function has no reference value, it does not go in the package.

## Four places where the conventional implementation is wrong

These are the reasons to use `riskval` rather than a snippet off the internet.

**1. The Hosmer–Lemeshow degrees of freedom depend on where the predictions
came from.** The familiar *G* − 2 is correct when the predictions come from a
logistic model fitted to the very data under test — two degrees of freedom go
on the fitted intercept and slope. In a *validation* exercise the predictions
are exogenous: a model estimated on a development sample, or a vendor model,
scored on fresh data. Nothing is estimated from the test sample, and the null
distribution is χ²*_G_*. `riskval` verifies both regimes by simulation:

| Predictions | Mean statistic (*G* = 10) | KS vs χ²₁₀ | KS vs χ²₈ |
| --- | --- | --- | --- |
| Exogenous | 9.86 | *p* = 0.58 | *p* = 9 × 10⁻¹⁸ |
| Fitted in-sample | 8.04 | *p* = 2 × 10⁻¹⁵ | *p* = 0.99 |

Using *G* − 2 on exogenous predictions *overstates* the evidence against the
model. `validate()` therefore defaults to *G*, and
`hosmer_lemeshow(..., df=...)` lets you say which you mean.

**2. The Jeffreys test is what the ECB prescribes, and it materially
outperforms the exact binomial test on small grades.** The binomial test is
conservative because the binomial is discrete. For a grade of 100 obligors at a
2% PD, at a nominal 5% level:

| Test | Realised size | Rejects from |
| --- | --- | --- |
| Exact binomial | 1.5% | 6 defaults |
| Jeffreys | 5.0% | 5 defaults |

**3. A one-sided supervisory test does not answer "is this model
calibrated".** Pillar 1 exists to catch insufficient capital, so the binomial
test a supervisor specifies is one-sided: *is the PD understated?* A model that
errs on the conservative side clears it while being badly miscalibrated. That
is not a prudential problem, but it is an economic one — capital held against
losses that will not happen, credit refused to borrowers who would have repaid.

`grade_level_report` therefore reports all three directions, and
`ValidationReport.summary()` shows the split:

```python
>>> import numpy as np, pandas as pd
>>> from riskval import validate
>>> rng = np.random.default_rng(1)
>>> pd_true = rng.uniform(0.005, 0.25, size=6000)
>>> y = rng.binomial(1, pd_true)
>>> grades = np.asarray(pd.cut(pd_true, bins=5, labels=list("ABCDE")))
>>>
>>> report = validate(y, pd_true * 1.6, grades=grades, n_boot=50, seed=0)
>>> frame = report.grade_report
>>> frame.attrs["rejection_rate_understated"]   # what a supervisor tests
0.0
>>> frame.attrs["rejection_rate_twosided"]      # whether it is calibrated
1.0

```

Every PD inflated by 60%, and the supervisory test flags nothing.

**4. The 1.06 IRB scaling factor.** Basel II and the EU CRR multiply IRB risk
weights by 1.06; the finalised Basel III framework removed it. Implementations
disagree and rarely say which they use. `riskval` defaults to 1.0, says so, and
takes `scaling_factor=1.06` for CRR figures.

`riskval` also documents where its tests *stop working*. Simulated under the
true model at the Basel 250-day window and 99% coverage, a nominal 5% Kupiec
test rejects correct models 9.1% of the time, while Christoffersen's
independence test rejects only 1.3%. Both figures are pinned in the test suite
and stated in the docstrings.

## The gap nobody tests: dispersion

Every test above asks whether the PD *level* is right — for the portfolio, or
grade by grade. None of them asks whether the predictions are *spread* right,
and that is what costs capital, because the Basel IRB risk weight is **concave
in PD**. Mean-preserving spread in the PDs therefore destroys risk-weighted
assets: an over-dispersed model reports less capital than the realised defaults
require, while clearing the level tests.

`calibration_slope` is the Cox (1958) recalibration slope, and it is in
`validate()` for exactly this reason. Take a portfolio and widen the spread of
its predictions on the logit scale, leaving the level alone:

| spread factor | Cox slope | one-sided supervisory rejection | capital deviation |
| --- | --- | --- | --- |
| 1.0 | 0.985 | 0.2 | −0.3% |
| 1.3 | 0.758 | 0.3 | −3.4% |
| 1.6 | 0.616 | **0.4** | −7.2% |
| 1.8 | 0.547 | **0.4** | −9.9% |
| 2.2 | 0.448 | **0.4** | −15.8% |

The capital shortfall doubles from −7.2% to −15.8% and the one-sided
supervisory rejection rate does not move at all. The slope tracks the damage;
the level test does not. Both columns are pinned by tests in
`tests/test_calibration.py`.

The complement holds too, which is why both belong in the battery: halve every
PD and the level tests reject while the slope correctly reports 1.0 — a level
failure is what the supervisory test is *for*.

### Use a materiality band, not a point null

A point null on 10 000 obligors rejects a slope of 0.97 — statistically real,
economically irrelevant. `materiality=m` tests `|b − 1| ≤ m` instead, a
minimal-effect test in the sense of Wellek (2010). Scored against the realised
capital deviation on 402 credit models:

| trigger rule | detects a >10% capital shortfall | false alarm |
| --- | --- | --- |
| slope test rejects at 5% (point null) | 99.5% | 76.9% |
| grade-level one-sided rejection rate > 20% | 31.6% | 38.5% |
| grade-level two-sided rate > 50% | 60.9% | 53.8% |
| **`materiality=0.15`, i.e. the interval clears [0.85, 1.15]** | **97.7%** | **26.9%** |

The rank correlation with the capital deviation is +0.80 for the slope and
+0.02 (p = 0.72) for the grade-level one-sided rate — the instrument in use
today is uncorrelated with the damage. Caveat: the false-alarm column rests on
26 cells whose capital was within 2% of correct, so read it as indicative.

```python
res = calibration_slope(y_true, y_prob, materiality=0.15)
res.statistic          # the slope
res.detail["band"]     # (0.85, 1.15)
res.reject             # materially over- or under-dispersed
```

## What it does not do

By design:

- **No model training.** `riskval` evaluates models; it does not build them.
- **No plotting.** It returns tables; you plot them.
- **No data downloading.** It is not a data source.
- **No proprietary dependencies.**

## Development

```bash
python -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"
pytest --cov=riskval --cov-report=term-missing
ruff check src tests
```

Test coverage is 100% and the floor is 85%. The suite includes the README's own
examples, so the quickstart above cannot go stale.

## Citing

If `riskval` contributed to published work, please cite it. See `CITATION.cff`.

## License

MIT. See `LICENSE`.
