Metadata-Version: 2.4
Name: causalatee
Version: 0.0.5
Summary: A toolkit for causality deep learning.
Author-email: Tim Hagen <tim.hagen@uni-kassel.de>
License: MIT License
        
        Copyright (c) 2026 Webis
        
        Permission is hereby granted, free of charge, to any person obtaining a copy
        of this software and associated documentation files (the "Software"), to deal
        in the Software without restriction, including without limitation the rights
        to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
        copies of the Software, and to permit persons to whom the Software is
        furnished to do so, subject to the following conditions:
        
        The above copyright notice and this permission notice shall be included in all
        copies or substantial portions of the Software.
        
        THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
        IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
        FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
        AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
        LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
        OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
        SOFTWARE.
        
Project-URL: Homepage, https://github.com/webis-de/causalatee
Project-URL: Bug Tracker, https://github.com/webis-de/causalatee/issues
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Operating System :: OS Independent
Classifier: Intended Audience :: Science/Research
Classifier: Topic :: Scientific/Engineering
Requires-Python: >=3.9
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: pandas[pyarrow]~=2.3
Requires-Dist: fastavro
Requires-Dist: openpyxl
Requires-Dist: datasets>=2.14
Requires-Dist: numpy>=1.24
Requires-Dist: scipy>=1.10
Provides-Extra: tests
Requires-Dist: bandit[toml]~=1.7; extra == "tests"
Requires-Dist: mypy~=1.5; extra == "tests"
Requires-Dist: pandas-stubs~=2.0; extra == "tests"
Requires-Dist: scipy-stubs>=1.10; extra == "tests"
Requires-Dist: pytest~=8.0; extra == "tests"
Requires-Dist: pytest-cov<8.0,>=5.0; extra == "tests"
Requires-Dist: ruff<0.15,>=0.9; extra == "tests"
Requires-Dist: types-pyyaml~=6.0; extra == "tests"
Requires-Dist: types-networkx~=3.0; extra == "tests"
Provides-Extra: huggingface
Requires-Dist: transformers>=4.40; extra == "huggingface"
Provides-Extra: baselines
Requires-Dist: spacy>=3.7; extra == "baselines"
Requires-Dist: networkx>=3.0; extra == "baselines"
Provides-Extra: mining
Requires-Dist: aiostream>=0.5; extra == "mining"
Provides-Extra: torch
Requires-Dist: torch>=2.0; extra == "torch"
Requires-Dist: torchmetrics>=1.0; extra == "torch"
Provides-Extra: docs
Requires-Dist: mkdocs~=1.6; extra == "docs"
Requires-Dist: mkdocs-material~=9.5; extra == "docs"
Requires-Dist: mkdocs-bibtex~=2.16; extra == "docs"
Requires-Dist: mkdocs-macros-plugin>=1.0; extra == "docs"
Requires-Dist: mkdocs-awesome-pages-plugin>=2.9; extra == "docs"
Requires-Dist: mkdocs-section-index>=0.3; extra == "docs"
Requires-Dist: mkdocs-jupyter>=0.24; extra == "docs"
Requires-Dist: mkdocstrings[python]<1.0,>=0.27; extra == "docs"
Requires-Dist: mike~=2.1; extra == "docs"
Requires-Dist: pypandoc_binary~=1.13; extra == "docs"
Dynamic: license-file

<p align="center">
   <img width=200px src="docs/assets/icon.png"/>
</p>

<p align="center">
  A Python library to simplify handling causality in natural language: extraction, training
  extraction models, unified datasets, and storing/querying results as causal graphs.
</p>

<p align="center">
  <a href="https://pypi.org/project/causalatee/"><img alt="PyPI" src="https://img.shields.io/pypi/v/causalatee"></a>
  <a href="https://causalatee.webis.de/"><img alt="Docs" src="https://img.shields.io/badge/docs-causalatee.webis.de-blue"></a>
  <img alt="Python" src="https://img.shields.io/pypi/pyversions/causalatee">
  <a href="https://github.com/webis-de/causalatee/actions/workflows/tests.yml"><img alt="Tests" src="https://github.com/webis-de/causalatee/actions/workflows/tests.yml/badge.svg"/></a>
  <a href="https://github.com/webis-de/causalatee/actions/workflows/linter.yml"><img alt="Linter" src="https://github.com/webis-de/causalatee/actions/workflows/linter.yml/badge.svg"/></a>
  <a href="https://codecov.io/gh/webis-de/causalatee"><img src="https://codecov.io/gh/webis-de/causalatee/graph/badge.svg?token=4jXcKrEUFc"/></a>
</p>

## What is this?

`causalatee` covers the full lifecycle of working with causality in text:

- **Datasets** — causality corpora, converted into one consistent HuggingFace-compatible schema
  across three standardized tasks:
  - **Causality Detection** — does a sentence express a causal relation at all?
  - **Causal Candidate Extraction** — which spans in a sentence are causes/effects?
  - **Causality Identification** — given two marked spans, does a causal relation hold between them?
- **`causalatee.models`** — batch-callable [`typing.Protocol`][protocol] interfaces
  (`Detection`, `CandidateExtraction`, `PairwiseIdentification`, `Identification`, `Extraction`)
  that any conforming model (rule-based, HuggingFace, or otherwise) satisfies with zero
  inheritance, plus `compose_extraction` to build an end-to-end extractor from the three
  sub-tasks.
- **`causalatee.graph`** — a typed `Graph`/`Node`/`Edge` interface with several interchangeable
  backends: an eager [CauseNet](https://causenet.org) loader, **CGF** (a compact,
  memory-mappable on-disk format), a [CausalBank](https://github.com/eecrazy/CausalBank)
  Cause-Effect Graph loader, and `SQLGraph`, a generic mutable graph you build yourself via
  `add_node`/`add_edge`.
- **`causalatee.mining`** — a streaming, concurrency-aware pipeline
  (`source -> flat_map -> filter -> map -> map -> reduce`) that turns a corpus of raw documents
  into an aggregated causal graph, without materializing the corpus in memory.
- **`causalatee.nn`** / **`causalatee.integrations`** — a biaffine span-grid extraction head, and
  ready-made HuggingFace `Pipeline` / PyTorch Lightning integrations for the three tasks above.

[protocol]: https://docs.python.org/3/library/typing.html#typing.Protocol

See **[causalatee.webis.de](https://causalatee.webis.de/)** for the full documentation,
including the dataset inventory, task/model reference, and runnable example notebooks.

## Installation

```bash
pip install causalatee
```

Optional extras pull in dependencies for specific pieces: `huggingface` (fine-tuning/inference
pipelines), `baselines` (dependency-parse baselines), `mining` (the corpus-mining pipeline), and
`docs` (building this documentation locally).

## Quick start

Every converted dataset is a standard HuggingFace dataset, indexed by task:

```python
from datasets import load_dataset

dataset = load_dataset("thagen/CausalNewsCorpus", "causality identification")
```

See the [Datasets](https://causalatee.webis.de/datasets/) page for the full list, and the
[Examples](https://causalatee.webis.de/examples/) notebooks for fine-tuning a model, building a
causal graph, and mining one from a raw corpus.

## Development

```bash
pip install -e ".[tests]"
ruff check .
mypy -p causalatee
pytest
```

## Building the Documentation

```bash
pip install -e ".[docs]"
mkdocs serve       # live-reloading local server
mkdocs build       # static site written to site/
```
