Metadata-Version: 2.4
Name: crosstabs
Version: 1.1.1
Summary: MCP package with 39 tools for statistics plus 23 headless research workflow tools
Project-URL: Homepage, https://crosstabs.com
Project-URL: Documentation, https://crosstabs.com
Project-URL: Source, https://pypi.org/project/crosstabs/#files
Project-URL: Support, https://crosstabs.com/support
Author-email: "crosstabs.com" <support@crosstabs.com>
License-Expression: MIT
License-File: LICENSE
Keywords: chi-square,contingency-tables,cramers-v,crosstabs,fisher-exact,mcp,model-context-protocol,odds-ratio,statistical-analysis,statistics
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering :: Information Analysis
Classifier: Topic :: Scientific/Engineering :: Mathematics
Requires-Python: >=3.10
Requires-Dist: fastmcp>=0.1.0
Requires-Dist: mcp>=1.0.0
Requires-Dist: numpy>=1.24.0
Requires-Dist: pandas>=2.0.0
Requires-Dist: scipy>=1.10.0
Requires-Dist: statsmodels>=0.14.0
Provides-Extra: dev
Requires-Dist: build>=1.2.2; extra == 'dev'
Requires-Dist: pip-audit>=2.9.0; extra == 'dev'
Requires-Dist: pytest-cov>=4.0.0; extra == 'dev'
Requires-Dist: pytest>=7.0.0; extra == 'dev'
Requires-Dist: tomli>=2.0.0; (python_version < '3.11') and extra == 'dev'
Description-Content-Type: text/markdown

# Crosstabs MCP Server

<!-- mcp-name: io.github.barangaroo/crosstabs -->

[![Python 3.10+](https://img.shields.io/badge/python-3.10+-blue.svg)](https://www.python.org/downloads/)
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)
[![MCP](https://img.shields.io/badge/MCP-Compatible-green.svg)](https://modelcontextprotocol.io)

One install provides two local MCP servers: **39 statistical tools** for focused contingency-table analysis and **23 headless research workflow tools** that take an agent from project creation and dataset import through tab books, coding, tracker repair, editable reports, and portable export. Individual statistical tools report whether their inference is exact, asymptotic, or simulated.

## Features

### Core Statistical Tests
| Test | Description |
|------|-------------|
| **Chi-square** | Pearson's chi-square test of independence |
| **G-test** | Likelihood-ratio alternative with an asymptotic p-value |
| **Fisher's exact** | Two-sided fixed-margin exact p-value for 2×2 integer counts |
| **McNemar's** | Exact two-sided binomial inference for fewer than 20 discordant pairs; continuity-corrected chi-square otherwise |

### Effect Sizes & Measures
| Measure | Use Case |
|---------|----------|
| **Cramér's V** | Effect size for any table size (with bias correction) |
| **Phi coefficient** | Effect size for 2×2 tables |
| **Odds ratio** | Association strength with a large-sample Woolf log interval |
| **Relative risk** | Risk comparison between groups |
| **Risk difference** | Absolute risk reduction |
| **Attributable risk** | Population-level impact |

### Ordinal Measures
| Measure | Description |
|---------|-------------|
| **Spearman's rho** | Rank correlation |
| **Kendall's tau** | Concordance measure |
| **Goodman-Kruskal gamma** | Ordinal association |
| **Somers' D** | Asymmetric ordinal measure |
| **Stuart's tau-c** | Rectangular table measure |

### Agreement & Reliability
| Measure | Description |
|---------|-------------|
| **Cohen's kappa** | Inter-rater agreement with an asymptotic normal interval |
| **Weighted kappa** | Linear/quadratic agreement with an asymptotic normal interval |

### Advanced Analysis
| Tool | Description |
|------|-------------|
| **CMH test** | Stratified analysis with a Robins-Breslow-Greenland pooled-OR interval |
| **Breslow-Day** | Test homogeneity of odds ratios |
| **Correspondence analysis** | Dimensionality reduction for tables |
| **Monte Carlo chi-square** | Fixed-margin simulated p-value estimate |
| **Power analysis** | Equal-group, two-sided normal approximation using Cohen's h |
| **Multiple comparisons** | Bonferroni and FDR corrections |

## Installation

### From PyPI (recommended)
```bash
pip install crosstabs
```

### Inspect the public source distribution
```bash
python -m pip download --no-deps --no-binary=:all: crosstabs
```

The development repository is currently private. PyPI publishes the package's
source archive; email
[support@crosstabs.com](mailto:support@crosstabs.com) for issue reports or
source-access questions.

## Quick Start

### Run the MCP Server
```bash
# Focused statistical calculators (Python 3.10+)
crosstabs

# End-to-end research workspace (Python 3.10+ and Node.js 20+)
crosstabs-headless
```

Or directly:
```bash
python -m crosstabs_mcp.server
```

### Configure an MCP client

Add to your `~/.claude/claude_desktop_config.json`:

```json
{
  "mcpServers": {
    "crosstabs_workspace": {
      "command": "crosstabs-headless"
    },
    "crosstabs_statistics": {
      "command": "crosstabs"
    }
  }
}
```

The headless server stores project state and generated artifacts under the local
application-data directory. Deterministic calculations stay on the machine.
`code_open_ends` is the only current operation that can use Vercel AI Gateway;
it requires an explicit external-processing approval in the tool call.

### Headless workflow tools

`create_project`, `import_dataset`, `profile_dataset`, `define_row_set`,
`define_banner`, `apply_filter`, `set_weight`, `run_table`, `run_tab_book`,
`compare_waves`, `code_open_ends`, `review_themes`, `approve_coding`,
`undo_change`, `replace_dataset`, `detect_schema_drift`, `repair_schema`,
`generate_report_pack`, `refresh_report_pack`, `export_project`,
`list_projects`, `inspect_project`, and `get_audit_history`.

Mutations use expected revisions and idempotency keys. Results include structured
warnings, evidence IDs, audit records, and MCP resources for generated files.

## Usage Examples

Once configured, Claude can perform statistical analysis:

### Chi-square Test
```
User: Test if there's an association between treatment and outcome:
      Treatment A: 50 success, 30 failure
      Treatment B: 20 success, 40 failure

Claude: [Uses chi_square_test with matrix [[50,30],[20,40]]]
        χ² = 11.67, p = 0.0006
        Cramér's V = 0.29 (small-medium effect)
        There is a significant association between treatment and outcome.
```

### Odds Ratio
```
User: Compare an adverse outcome between exposure groups:
      Exposed: 30 outcome-present, 70 outcome-absent
      Unexposed: 15 outcome-present, 85 outcome-absent

Claude: [Uses odds_ratio with matrix [[30,70],[15,85]]]
        OR = 2.43 (95% CI: 1.21-4.87)
        Exposure is associated with 143% higher odds of the outcome.
```

The epidemiology tools (`odds_ratio`, `relative_risk`, `risk_difference`, and
`attributable_risk`) use one explicit orientation:
`[[exposed outcome+, exposed outcome-], [unexposed outcome+, unexposed outcome-]]`.
They return machine-readable zero/infinite/undefined states rather than silently
continuity-correcting point estimates. Attributable/prevented fractions require
causal identification assumptions; an association alone does not establish the
counterfactual effect of removing an exposure.

### Fisher's Exact Test
```
User: I have a small sample: [[3,1],[1,5]]. Is it significant?

Claude: [Uses fishers_exact with the matrix]
        p = 0.190476 (two-tailed exact)
        Not statistically significant at α=0.05.
```

## Available Tools

| Tool Name | Description |
|-----------|-------------|
| `chi_square_test` | Chi-square test of independence |
| `g_test` | G-test (likelihood ratio) |
| `fishers_exact` | Fisher's exact test (2×2) |
| `mcnemar_test` | McNemar's test for paired data |
| `odds_ratio` | Odds ratio with CI |
| `relative_risk` | Relative risk with CI |
| `risk_difference` | Risk difference with CI |
| `cramers_v` | Cramér's V effect size |
| `phi_coefficient` | Phi for 2×2 tables |
| `cohens_kappa` | Cohen's kappa |
| `weighted_kappa` | Weighted kappa |
| `spearmans_rho` | Spearman's rank correlation |
| `kendalls_tau` | Kendall's tau-b |
| `goodman_kruskal_gamma` | Gamma coefficient |
| `somers_d` | Somers' D |
| `tau_c` | Stuart's tau-c |
| `cmh_test` | Cochran-Mantel-Haenszel |
| `breslow_day_test` | Breslow-Day test |
| `linear_trend_test` | Linear-by-linear association |
| `correspondence_analysis` | Correspondence analysis |
| `monte_carlo_chi_square` | Fixed-margin Monte Carlo p-value estimate |
| `power_analysis` | Two-sided normal-approximation power/sample size using Cohen's h |
| `bonferroni_correction` | Bonferroni p-value adjustment |
| `fdr_correction` | Benjamini-Hochberg FDR |
| `standardized_residuals` | Cell residuals |
| `post_hoc_chi_square` | Post-hoc chi-square decomposition |
| `proportion_ci` | Confidence interval for proportion |
| `check_assumptions` | Validate chi-square assumptions |
| `recommend_test` | Method suggestions with assumption caveats |
| `mosaic_plot_data` | Data for mosaic visualization |
| `stacked_bar_data` | Data for stacked bar chart |
| `attributable_risk` | Attributable risk measures |
| `chi_square_yates` | Yates' continuity correction |
| `effect_size` | Multiple contingency-table effect sizes |
| `lambda_coefficient` | Goodman–Kruskal lambda |
| `uncertainty_coefficient` | Theil's uncertainty coefficient |
| `detect_outliers` | Outlier detection |
| `crosstab_from_data` | Build table from raw data |
| `crosstab_from_csv` | Build table from CSV |

## Development

### Run Tests
```bash
pip install -e ".[dev]"
pytest tests/ -v
```

### Project Structure
```
mcp-server-python/
├── crosstabs_mcp/
│   ├── __init__.py
│   ├── headless_launcher.py # Node version check and bundled-server launcher
│   ├── headless-mcp.mjs     # Local-first 23-tool workflow MCP server
│   ├── server.py          # Main MCP server
│   └── advanced_stats.py  # Compatibility imports; math lives in server.py
├── tests/
│   ├── test_statistics.py       # Statistical behavior tests
│   └── test_reference_parity.py # Public SciPy reference parity
├── scripts/                    # Distribution verification
├── LICENSE
├── pyproject.toml
├── uv.lock
└── README.md
```

## Requirements

- Python 3.10+
- mcp >= 1.0.0
- fastmcp >= 0.1.0
- numpy >= 1.24.0
- scipy >= 1.10.0
- pandas >= 2.0.0
- statsmodels >= 0.14.0

## License

MIT License - see [LICENSE](LICENSE) for details.

## Contributing

The development repository is currently private. Send corrections and proposed
changes to [support@crosstabs.com](mailto:support@crosstabs.com).

## Links

- [PyPI package and source archive](https://pypi.org/project/crosstabs/#files)
- [Web Application](https://www.crosstabs.com)
- [Support](https://www.crosstabs.com/support)
- [MCP Documentation](https://modelcontextprotocol.io)
