Metadata-Version: 2.5
Name: mescla
Version: 0.1.0
Summary: End-Member Mixing Analysis (EMMA) and water mixing-ratio methods for hydrochemistry
Project-URL: Homepage, https://github.com/amphos21-consulting/mescla
Project-URL: Repository, https://github.com/amphos21-consulting/mescla
Project-URL: Issues, https://github.com/amphos21-consulting/mescla/issues
Project-URL: Documentation, https://amphos21-consulting.github.io/mescla
Author: Albert Nardi, Amphos 21 Consulting S.L.
License-Expression: MIT
License-File: LICENSE
Keywords: EMMA,end-member mixing analysis,geochemistry,hydrochemistry,hydrogeology,hydrograph separation,mixing ratios,tracers
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Science/Research
Classifier: Natural Language :: English
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Topic :: Scientific/Engineering
Classifier: Topic :: Scientific/Engineering :: Chemistry
Classifier: Topic :: Scientific/Engineering :: Hydrology
Requires-Python: >=3.11
Requires-Dist: numpy>=1.24
Requires-Dist: pandas>=2.0
Requires-Dist: scipy>=1.10
Provides-Extra: dev
Requires-Dist: matplotlib>=3.7; extra == 'dev'
Requires-Dist: mypy>=1.8; extra == 'dev'
Requires-Dist: pytest-cov>=4.1; extra == 'dev'
Requires-Dist: pytest>=7.4; extra == 'dev'
Requires-Dist: ruff>=0.5; extra == 'dev'
Provides-Extra: docs
Requires-Dist: furo>=2024.1; extra == 'docs'
Requires-Dist: ipykernel>=6.25; extra == 'docs'
Requires-Dist: matplotlib>=3.7; extra == 'docs'
Requires-Dist: myst-nb>=1.1; extra == 'docs'
Requires-Dist: sphinx-copybutton>=0.5; extra == 'docs'
Requires-Dist: sphinx-design>=0.5; extra == 'docs'
Requires-Dist: sphinx>=7.2; extra == 'docs'
Requires-Dist: wqchartpy>=0.1.10; extra == 'docs'
Provides-Extra: notebooks
Requires-Dist: ipykernel>=6.25; extra == 'notebooks'
Requires-Dist: jupyterlab>=4.0; extra == 'notebooks'
Requires-Dist: matplotlib>=3.7; extra == 'notebooks'
Requires-Dist: nbconvert>=7.0; extra == 'notebooks'
Requires-Dist: wqchartpy>=0.1.10; extra == 'notebooks'
Provides-Extra: plotting
Requires-Dist: matplotlib>=3.7; extra == 'plotting'
Description-Content-Type: text/markdown

<picture>
  <source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/amphos21-consulting/mescla/main/docs/_static/logo-lockup-dark.svg">
  <img src="https://raw.githubusercontent.com/amphos21-consulting/mescla/main/docs/_static/logo-lockup.svg" alt="Mescla" width="300">
</picture>

[![CI](https://github.com/amphos21-consulting/mescla/actions/workflows/ci.yml/badge.svg)](https://github.com/amphos21-consulting/mescla/actions/workflows/ci.yml)
[![Docs](https://github.com/amphos21-consulting/mescla/actions/workflows/docs.yml/badge.svg)](https://amphos21-consulting.github.io/mescla)
[![PyPI](https://img.shields.io/pypi/v/mescla.svg)](https://pypi.org/project/mescla/)
[![Python](https://img.shields.io/pypi/pyversions/mescla.svg)](https://pypi.org/project/mescla/)
[![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE)

A Python library for **End-Member Mixing Analysis (EMMA)** and **water mixing-ratio**
calculations in hydrochemistry, with a test suite that reproduces published results and a
set of teaching notebooks.

Two method families, used together:

| | **EMMA** | **Mixing ratios** |
|---|---|---|
| Answers | *How many* end-members are needed, and *which* species behave as a mixture | *In what proportions* those end-members mix in each sample |
| Machinery | Eigen-analysis of standardised chemistry; projection into U-space | Constrained least squares, or maximum likelihood when end-members are uncertain |
| Cannot | give you mixing ratios | tell you whether your conceptual model is right |
| Key refs | Christophersen & Hooper (1992); Hooper (2003) | Carrera et al. (2004) |

---

## Install

```bash
pip install mescla
```

To work on the library itself — the project uses [uv](https://docs.astral.sh/uv/), but
plain `pip` works equally well:

```bash
uv venv && uv pip install -e ".[dev,notebooks]"
```

Runtime dependencies are NumPy, SciPy and pandas. Matplotlib is needed only for
`mescla.plotting`. Tested on Python 3.11, 3.12, 3.13 and 3.14.

Full documentation, including the ten tutorials rendered with their output:
**<https://amphos21-consulting.github.io/mescla>**

## Tasks

Common jobs are [`just`](https://just.systems) recipes. `just` on its own lists them.

```bash
just setup       # create .venv and install everything (once per clone)
just lab         # launch JupyterLab on the notebooks
just fast        # the quick test loop (~14 s; skips the slow sweeps)
just check       # the pre-commit gate: lint + full suite
just docs        # build the documentation into docs/_build/html
```

The ones worth knowing about:

| recipe | what it is for |
|---|---|
| `just repro` | reproduce the published results of Carrera (2004) and Tubau (2014) — run this before believing a change to the estimators is harmless |
| `just notebooks` | execute all ten notebooks from a fresh kernel, as CI does |
| `just nb-clean` | strip notebook outputs; they are committed cleared so the repo stays diffable |
| `just smoke` | install the built wheel in a clean environment and drive a workflow from outside the repo |
| `just pyversions` | run the suite on every Python version the package claims to support |
| `just release-check` | audit everything that must be true before the repository is made public |

`just lint` treats ruff as blocking and mypy as advisory; `just types` shows the full
mypy report.

## Sixty seconds

```python
import mescla as am
from mescla.datasets import make_mixture

data = make_mixture(n_endmembers=3, n_species=5, n_samples=60, seed=0)

model = am.EMMA().fit(data.samples)     # we do not tell it the answer
print(model.n_endmembers)               # 3

result = am.mixing_ratios(data.endmembers, data.samples)
result.ratios_frame().head()
```

On your own data:

```python
import pandas as pd
import mescla as am

samples = am.WaterChemistry.from_frame(pd.read_csv("samples.csv", index_col=0))
endmembers = am.EndMembers.from_frame(pd.read_csv("endmembers.csv", index_col=0))

am.charge_balance_report(samples)                 # screen before you model
model = am.EMMA().fit(samples)                    # how many sources does the data need?
print(model.summary())
print(model.hull_fraction(endmembers))            # do yours bound the samples?

result = am.mixing_ratios(endmembers, samples, sigma=0.05 * samples.data)
result.ratios_frame().to_csv("mixing_ratios.csv")
```

## Notebooks

Run `jupyter lab notebooks/`. Each is narrative-first and ends with a self-contained
snippet to paste into your own project.

| | | |
|---|---|---|
| `00_quickstart` | synthetic | the whole pipeline in ~15 lines |
| `01_data_and_qaqc` | real | units, charge balance, seeing the chemistry first |
| `02_emma_rank_and_endmembers` | real | standardising, rank criteria compared, Hooper diagnostics, hull screening, archetypes |
| `03_mixing_ratios_least_squares` | synthetic + real | the closed forms, weighting, what a negative ratio means |
| `04_mix_maximum_likelihood` | published tables | MIX vs least squares as samples accumulate; **assigning sigma** |
| `05_uncertainty` | synthetic | propagation, Monte Carlo, resampling, identifiability |
| `06_reactions_from_residuals` | synthetic | reading dissolution and redox out of the misfit |
| `07_case_study_end_to_end` | real | a complete analysis with its limitations |
| `08_isotopes` | synthetic | water vs solute isotopes, and the trap in the second |
| `09_hubbard_brook_precipitation` | real | 50 years of acid rain: when the end-members do *not* hold still |

Figures appear where they earn their place: bivariate tracer plots in `01`, loadings in
`02`, uncertainty intervals in `05`, the reaction chart in `06`.

Notebooks are committed with outputs cleared and executed in CI, so they cannot rot.

## What is implemented

**`mescla.emma`** — correlation and covariance PCA; four rank criteria side by side
(rule of one, cumulative variance, broken stick, Horn's parallel analysis); U-space
projection and subspace distance; Hooper (2003) relative bias and RRMSE, a formal
residual-structure test, and cross-site projection; convex-hull screening of candidate
end-members; and **archetypal analysis** (Cutler & Breiman, 1994), which proposes vertices
when you have no candidate waters at all.

**`mescla.mixing`** — constrained weighted least squares (Carrera eq. 9, with the
active-set non-negativity step his convexity argument licenses); closed-form two- and
three-component formulas; the U-space geometric solve; and **MIX**, the maximum-likelihood
estimator for uncertain end-members.

**`mescla.uncertainty`** — Genereux (1998) propagation, Monte Carlo respecting the
simplex constraints, jackknife and bootstrap over end-member replicates, and an
identifiability report (condition number, end-member separation, hull fraction, sigma
sweep).

**`mescla.reactions`** — measured-versus-predicted comparison, source/sink
classification with suggested processes, and `fit_conservative`, which fits the ratios on
the tracers you trust so the reacting species can be judged against them.

**`mescla.isotopes`** — the distinction between isotopes of the water molecule (which
mix linearly) and isotope ratios of a solute (which do not, and must be weighted by their
carrier element). `linearize` turns the second kind into a quantity that mixes linearly, so
the rest of the library works on it unchanged; the estimators warn if you forget.

**Archetypal analysis** deserves its own paragraph, because it is the one unsupervised
method in the library that did not come from the mixing literature. It writes every sample
as a convex combination of *k* archetypes lying on the convex hull — non-negative weights
summing to one, which is the mixing geometry imposed by the objective rather than checked
afterwards. That is why it, and not a clustering algorithm, is the right unsupervised tool
here: a cluster is a crowded region of the mixing polytope, not a source. It comes with five
cautions, each of them enforced or reported rather than merely documented — see
`mescla.emma.archetypes` and notebook 02. The first is the one that matters most:
**archetypal weights are not mixing ratios**, and `ArchetypeResult.ratios` raises rather
than let the confusion pass.

**`mescla.prep`** — unit conversion with a major-ion table, charge balance
(`convention="sum"`, the Freeze & Cherry / PHREEQC factor of 100, or `"mean"` for the
APHA 1030E factor of 200 — the two differ by exactly 2x, and the ±5% rule is quoted for
both), EC
cross-check, a `Standardizer` that makes it impossible to standardise end-members with the
wrong statistics, and `prep.missing` for ragged datasets: what each species costs you in
samples, and the complete sub-matrix that retains the most data.

**`mescla.plotting`** — U-space mixing diagram, scree, eigenvector loadings, residual
small multiples, measured-versus-predicted, bivariate mixing line, stacked ratios, ternary,
mixing ratios **with their uncertainty intervals**, a diverging source/sink chart of
reaction departures. The palette is validated for colour-vision safety rather than chosen
by eye, and identity is carried by marker shape and direct labels as well as colour.

Classical water-type diagrams (Piper, Schoeller, Stiff, Durov and the rest) are out of
scope: [WQChartPy](https://github.com/jyangfsu/WQChartPy) already covers them, takes a
plain DataFrame and is a better tool for that job than anything this library would ship.

## Validation

```bash
pytest                      # 319 tests
pytest -m "not slow"        # quicker
ruff check src tests
```

Two reproduction suites assert the *published claims* of the source papers, which is what
makes the library trustworthy rather than merely functional.

`tests/test_reproduce_carrera2004.py`:

- least-squares correlations fall in the published 0.97–0.98 (low end-member noise) and
  0.93–0.94 (high noise) bands;
- least squares does **not** improve with sample count, while MIX does;
- MIX beats least squares at every sample count and noise level tested;
- MIX is insensitive to end-member noise where least squares is not;
- the improvement index lands in the published 3–8 range.

`tests/test_reproduce_tubau2014.py` checks the published Besòs end-member table and sigma
scheme, and reproduces a *qualitative* claim as a number: the paper says its two dry-period
end-members are "distinguished mainly by high ammonium and low calcium and magnesium", and
our identifiability diagnostics independently find that on the four conservative species
alone those end-members have condition number 159 and relative separation 0.105 — a failure
by our own thresholds — which adding NH₄, Ca and Mg resolves.

The MIX implementation is also cross-checked against the paper's own formulation:
`profile_objective` computes the eliminated objective of equation 21, and the test suite
asserts it agrees with what the solver minimises.

## Data

Provenance for everything bundled is in `src/mescla/datasets/data/SOURCES.md`.

- **Synthetic** — `make_mixture`, plus `carrera_application1` and `carrera_application2`
  reproducing the paper's two test problems. These carry the ground truth, which is what
  makes validation possible at all.
- **`load_tubau_besos()`** — the three published Besòs River end-members and, more usefully,
  the published standard-deviation scheme from Tubau et al. (2014).
- **`load_grafton_nh()`** — 130 real major-ion analyses (66 stream, 64 well) from the
  county containing Hubbard Brook, via the Water Quality Portal. USGS, public domain.
- **`fetch_water_quality_portal()`** — live, cached, reproducible queries for your own sites.
- **`load_hubbard_brook()`** — 600 complete months of bulk precipitation chemistry at
  Hubbard Brook Watershed 6, 1963–2014, from the Environmental Data Initiative (CC-BY,
  attribution required). Nine species including H⁺ derived from the published pH. Pass
  `annual=True` for volume-weighted water-year means. This is the record where you can
  *watch* an end-member move: non-sea-salt sulphate falls roughly fivefold across it.

## Notes on use

A few things the library will not do for you, drawn from the papers it implements:

- **Units are mg/L by default, and tracked per species.** A table can hold mg/L beside
  µS/cm and permil, and say so. Mixing does not depend on the label; conversions and
  charge balance do. `to_meq_per_l` **refuses** a species it cannot convert rather than
  passing it through silently — pass `on_unknown="skip"` or `"drop"` to choose. Inspect
  the factors with `conversion_factors(table)`.
- **Mixing is linear in concentration.** Never log-transform concentrations; never mix
  ratios of species. δ¹⁸O and δ²H of water mix linearly, but isotope ratios of a *solute*
  (δ¹³C-DIC, ⁸⁷Sr/⁸⁶Sr) must be weighted by the carrier element — use `linearize()`.
  Getting this wrong is not a small bias: with a 20× contrast in the carrier, a true
  fraction of 0.75 comes back as 0.13.
- **A non-detect is not a measurement.** It says the concentration lies in [0, limit].
  Carry it as `censored` rather than substituting half the limit, which is biasing and in
  our tests was worse than doing nothing at all.
- **A conservative tracer is conservative in *this* system.** SO₄ is fine in an oxic
  aquifer and useless where sulphate reduction occurs.
- **Assigning sigma is a modelling decision that changes the answer.** Sweep it and report
  the sweep.
- **Fit the ratios on conservative tracers before diagnosing reactions.** Otherwise least
  squares absorbs the reaction into the ratios and blames an innocent species.
- **A charge balance error is meaningless without its convention.** The two in circulation
  differ by a factor of two, and the ±5% rule is quoted for both. `charge_balance_report`
  defaults to the Freeze & Cherry / PHREEQC form and records which it used.
- **The hardest assumption is that the end-members held still.**
- **An unsupervised method proposes; it does not discover.** Archetypal analysis will return
  *k* vertices from any data you give it, whether or not they are waters that exist. Run it
  after the conceptual model is written down, as a check on it -- which is how notebook 07
  uses it, and where it agrees about two sources out of three and points at an outlier with
  the third.

## References

Carrera, J., Vázquez-Suñé, E., Castillo, O. and Sánchez-Vila, X. (2004). A methodology to
compute mixing ratios with uncertain end-members. *Water Resources Research* **40**,
W12101. doi:10.1029/2003WR002263

Cutler, A. and Breiman, L. (1994). Archetypal analysis. *Technometrics* **36**(4), 338-347.

Christophersen, N. and Hooper, R. (1992). Multivariate analysis of stream water chemical
data: the use of principal components analysis for the end-member mixing problem.
*Water Resources Research* **28**(1), 99–107. doi:10.1029/91WR02518

Genereux, D. (1998). Quantifying uncertainty in tracer-based hydrograph separations.
*Water Resources Research* **34**(4), 915–919. doi:10.1029/98WR00010

Hooper, R. (2003). Diagnostic tools for mixing models of stream water chemistry.
*Water Resources Research* **39**(3), 1055. doi:10.1029/2002WR001528

Tubau, I., Vàzquez-Suñé, E., Jurado, A. and Carrera, J. (2014). Using EMMA and MIX analysis
to assess mixing ratios and to identify hydrochemical reactions in groundwater.
*Science of the Total Environment* **470–471**, 1120–1131.
doi:10.1016/j.scitotenv.2013.10.121

Likens, G. (2016). Chemistry of Bulk Precipitation at Hubbard Brook Experimental Forest,
Watershed 6, 1963 – present, ver 9. *Environmental Data Initiative*.
doi:10.6073/pasta/8d2d88dc718b6c5a2183cd88aae26fb1 (CC-BY)

Read, E.K. et al. (2017). Water quality data for national-scale aquatic research: the Water
Quality Portal. *Water Resources Research* **53**, 1735–1745. doi:10.1002/2016WR019993

Further background on tracer selection and the wider literature is in the guides under
[`docs/guides/`](docs/guides/), and in the module docstrings, which cite their sources.

## Acknowledgements

Mescla was developed at **Amphos 21 Consulting S.L.**, which funded the work.

The methods are not ours: they are due to Christophersen and Hooper, to Hooper, and to
the Barcelona hydrogeology group of Carrera, Vázquez-Suñé, Tubau and colleagues. This
library is an implementation and a test of their published results, and it cites them
throughout.

## Contributing

[`CONTRIBUTING.md`](CONTRIBUTING.md) has the setup, the checks to run before opening a
pull request, and the few conventions specific to this repository. `CHANGELOG.md`
records what changed in each release.

## Licence and attribution

The **code** is MIT licensed (see `LICENSE`).

The **bundled data is not** — each file carries its own terms, recorded in full in
[`src/mescla/datasets/data/SOURCES.md`](src/mescla/datasets/data/SOURCES.md):

| dataset | terms |
|---|---|
| `grafton_nh_major_ions.csv` | U.S. Geological Survey — public domain |
| `hubbard_brook_ws6_precipitation.csv` | Hubbard Brook Ecosystem Study — **CC-BY 4.0** |
| `tubau2014_besos_endmembers.csv` | short factual tables from a published paper, reproduced with attribution |

**CC-BY makes attribution a condition of use, not a courtesy.** If you publish a figure
or a result derived from the Hubbard Brook record, the citation in `SOURCES.md` must
travel with it.

The reference papers themselves are publisher copyright and are **not** distributed with
this repository. Keep your own copies in `papers/`, which is gitignored; every module
that implements a method cites its source.

If Mescla is useful in your work, `CITATION.cff` has the software citation and the
four methods papers it implements.
