Metadata-Version: 2.4
Name: gallop-pds
Version: 0.4.0
Summary: Product data science checks: power from priors, SRM, CUPED, always-valid inference, empirical Bayes shrinkage, out-of-time model validation, mix-versus-rate decomposition, and the prior store they read.
Author: 0trm
License-Expression: MIT
Project-URL: Homepage, https://0trm.github.io/gallop/
Project-URL: Repository, https://github.com/0trm/gallop
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Science/Research
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering :: Information Analysis
Requires-Python: >=3.11
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: numpy>=1.26
Requires-Dist: pandas>=2.0
Requires-Dist: scipy>=1.11
Requires-Dist: rich>=14
Requires-Dist: mcp<3,>=2.2
Provides-Extra: dev
Requires-Dist: pytest>=8; extra == "dev"
Requires-Dist: ruff<0.17,>=0.16; extra == "dev"
Requires-Dist: markdown-it-py>=3; extra == "dev"
Dynamic: license-file

# gallop

<br>
<p align="center">
  <img src="site/assets/mark.svg" width="300" alt="Three riders carried on one galloping horse">
</p>

<p align="center">
  <a href="#install">install</a> &middot; <a href="#quick-start">quick start</a> &middot; <a href="https://0trm.github.io/gallop/skills/">docs</a>
</p>

<p align="center">
  <a href="LICENSE"><img src="https://img.shields.io/badge/license-MIT-999999?labelColor=555555" alt="MIT license" /></a>
  <a href="https://github.com/0trm/gallop/releases/latest"><img src="https://img.shields.io/github/v/release/0trm/gallop?label=release&labelColor=555555&color=999999" alt="latest stable release" /></a>
  <a href="https://0trm.github.io/gallop/"><img src="https://img.shields.io/badge/docs-github%20pages-999999?labelColor=555555&logo=github&logoColor=white" alt="documentation site on GitHub Pages" /></a>
</p>

<p align="center">
  <strong>Skills that make your coding agent think like a product data scientist.</strong><br>
  They decide if a question deserves an analysis, pick the method, and check the result.
</p>

## Install

```bash
# Claude Code
/plugin marketplace add 0trm/gallop
/plugin install gallop@gallop

# Any other agent, or none: a skill is a directory of markdown
git clone https://github.com/0trm/gallop
cp -r gallop/skills/reading-experiments .claude/skills/

# The package; the import name is gallop
pip install gallop-pds
```

## Quick start

One command, a synthetic dataset, and every check once:

```bash
python3 -m gallop.examples.quickstart
```

The full loop, from a question arriving to the prior store changing on disk, is `python3 examples/end-to-end/run_loop.py`.

| Bucket | Asks | Hands back |
|---|---|---|
| Description | What happened? | A hypothesis |
| Causation | Did this change cause that? | An effect size |
| Prediction | What will happen? Who gets what? | A forecast, a ranking, an allocation |

## Map

A question enters at the left and leaves as a decision. Measurement is a foundation because what ships changes the data. Theory is a ceiling because what you learn has to outlive the test that produced it.

<picture>
  <source media="(prefers-color-scheme: dark)" srcset="site/figures/skills-map-dark.svg">
  <img src="site/figures/skills-map-light.svg" width="100%" alt="The eight skills placed on the method map: routing at the entry, a can-you-randomise diamond, experimentation and causal inference to the right, exploratory analytics and statistical modeling below the path, the measurement framework as the floor and the theory layer as the ceiling">
</picture>

> [Read the full map](https://0trm.github.io/gallop/map/)

## Skills

<!-- skills-table:begin (generated by site/build.py; do not edit) -->
| Skill | What it decides | Reach for it when |
|---|---|---|
| [`routing-questions`](skills/routing-questions/SKILL.md) | Whether this becomes work at all, and which skill it becomes | a product, analytics, or experimentation request first arrives, when someone asks for a deep dive or a dashboard, or before opening a query editor on any question about impact, lift, or whether something worked |
| [`defining-metrics`](skills/defining-metrics/SKILL.md) | A metric turned into a computation, a source of truth, a registry entry, and a statement of how it will be gamed | defining a north-star or guardrail metric, when two dashboards disagree on the same number, when arbitrating between conflicting metric definitions, or when a readout depends on a metric nobody has validated |
| [`sizing-opportunities`](skills/sizing-opportunities/SKILL.md) | A what-happened question turned into a localised, sized hypothesis, with the floor checked first and the gap never quoted as the prize | a metric moved and someone asks what happened, when asked for a deep dive, a funnel or segment analysis, a root cause, or an opportunity size before a roadmap commitment, or when an observed gap between two groups is about to be quoted as the value of closing it |
| [`designing-experiments`](skills/designing-experiments/SKILL.md) | The four choices that cannot be repaired after launch, with the MDE from the prior store | planning, powering, or pre-registering an experiment, when deciding whether a question is testable at the available traffic, or when a feature is about to ship without a flag |
| [`reading-experiments`](skills/reading-experiments/SKILL.md) | Whether the result is a result: SRM, exposure, the sequential bound, CUPED, shrinkage | analysing or reviewing A/B test results, when a test looks like a winner, when someone reports a lift, or when deciding whether to ship on an experiment readout |
| [`choosing-causal-designs`](skills/choosing-causal-designs/SKILL.md) | The method that matches how assignment happened, and the exit that says there is no comparison group | measuring the impact of something already rolled out, a launch, a migration, a pricing change, or a campaign that reached everyone at once |
| [`automating-decisions`](skills/automating-decisions/SKILL.md) | Whether a forecast or a repeated decision belongs to a model, validated out of time, and the holdout that measures its impact | someone asks for a churn, propensity, LTV, scoring, forecasting, uplift, recommendation or allocation model, when a model's offline accuracy is offered as evidence that something worked, or when deciding who gets an offer, a discount or an intervention |
| [`writing-readouts`](skills/writing-readouts/SKILL.md) | The decision rule first, the result last; the belief filed where the next question starts | a test finishes, when documenting a shipped or killed decision, when writing up a null result or a rollback, or when a question needs an entry someone can find in a year |
<!-- skills-table:end -->

## License

[MIT](LICENSE). What it covers, what gallop does with your data, and how to contribute: [docs/model.md](docs/model.md).
