Metadata-Version: 2.4
Name: tablassert
Version: 19.5.1
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Development Status :: 5 - Production/Stable
Classifier: Intended Audience :: Science/Research
Classifier: Intended Audience :: Healthcare Industry
Classifier: Intended Audience :: Developers
Classifier: Topic :: Scientific/Engineering :: Bio-Informatics
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Scientific/Engineering :: Information Analysis
Classifier: Topic :: Database
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Programming Language :: Python :: Implementation :: CPython
Classifier: Programming Language :: Rust
Classifier: Framework :: Pydantic :: 2
Classifier: Operating System :: POSIX :: Linux
Classifier: Operating System :: MacOS :: MacOS X
Classifier: Environment :: Console
Requires-Dist: biolink-model>=4.4.5
Requires-Dist: polars>=2.0.0rc2
Requires-Dist: pyarrow>=21.0.0
Requires-Dist: rapidfuzz>=3.14.3
Requires-Dist: pydantic>=2.12.5
Requires-Dist: pyyaml>=6.0.3
Requires-Dist: fastexcel>=0.21.0
Requires-Dist: smolagents>=1.26.0 ; extra == 'agent'
Requires-Dist: litellm>=1.93.0 ; extra == 'agent'
Requires-Dist: aria2==0.0.1b0 ; extra == 'aria2'
Requires-Dist: cyclopts>=1.0.0 ; extra == 'cli'
Requires-Dist: rich>=13.0.0 ; extra == 'cli'
Requires-Dist: datasets>=3.0.0 ; extra == 'distill'
Requires-Dist: loguru>=0.7.3 ; extra == 'log'
Requires-Dist: dspy>=3.2.1 ; extra == 'optimize'
Requires-Dist: scikit-learn>=1.8.0 ; extra == 'qc'
Requires-Dist: sentence-transformers>=5.3.0 ; extra == 'qc'
Requires-Dist: polars[rtcompat]>=2.0.0rc2 ; extra == 'rt'
Provides-Extra: agent
Provides-Extra: aria2
Provides-Extra: cli
Provides-Extra: distill
Provides-Extra: log
Provides-Extra: optimize
Provides-Extra: qc
Provides-Extra: rt
License-File: LICENSE
Summary: Tablassert turns biomedical spreadsheets (Excel, CSV, TSV) into knowledge graphs ready for NCATS Translator. Declare how your columns map to subject–predicate–object statements in YAML; Tablassert resolves free text to standard CURIEs, attaches provenance and statistical annotations, and emits KGX-compliant nodes and edges.
Keywords: knowledge graph,kgx,ncats translator,biolink,bioinformatics,biomedical data,entity resolution,tabular data,yaml configuration,declarative pipeline,quality control,llm agent
Author-email: Skye Lane Goetz <sgoetz@isbscience.org>
License-Expression: Apache-2.0
Requires-Python: >=3.11
Description-Content-Type: text/markdown; charset=UTF-8; variant=GFM
Project-URL: Documentation, https://skyeav.github.io/Tablassert/
Project-URL: Homepage, https://github.com/SkyeAv/Tablassert
Project-URL: Source, https://github.com/SkyeAv/Tablassert

# Tablassert

[![PyPI version](https://badge.fury.io/py/tablassert.svg)](https://badge.fury.io/py/tablassert)
[![CI](https://github.com/SkyeAv/Tablassert/actions/workflows/ci.yml/badge.svg)](https://github.com/SkyeAv/Tablassert/actions/workflows/ci.yml)
[![Docs](https://img.shields.io/github/deployments/SkyeAv/Tablassert/github-pages?label=docs)](https://skyeav.github.io/Tablassert/)
[![License](https://img.shields.io/badge/License-Apache%202.0-blue.svg)](https://opensource.org/licenses/Apache-2.0)

> Extract knowledge assertions from tabular data into NCATS Translator-compliant KGX NDJSON,
> declaratively, with entity resolution built in and optional quality control.

Tablassert turns biomedical spreadsheets (Excel, CSV, TSV) into knowledge graphs ready for NCATS
Translator. Declare how your columns map to subject-predicate-object statements in YAML; Tablassert
resolves free text to standard CURIEs, attaches provenance and statistical annotations, and emits
KGX-compliant nodes and edges.

**[Full Documentation](https://skyeav.github.io/Tablassert/)**: installation guides, tutorial,
configuration reference, and API docs.

## Statement of need

Biomedical knowledge lives in spreadsheets: association tables, assay results, curated gene-disease
lists. Getting those rows into an NCATS Translator knowledge graph means writing an ingest that maps
columns to Biolink statements, resolves free text ("TP53", "lung cancer") to standard CURIEs, and
records provenance, terms of use, and statistical annotations. Today that ingest is Python code
maintained per source, and entity resolution usually means calling a hosted name-resolution service
at build time.

Tablassert replaces the per-source code with a declarative YAML mapping plus an embedded, offline
entity-resolution database built from RENCI BABEL exports, so a build is reproducible, auditable,
and network-free. It is aimed at Translator ingest authors and knowledge-graph data engineers who
hold tabular biomedical sources, and at bioinformatics groups that want a validated KGX output
without writing a transform pipeline. The configuration model and the Biolink/KGX contracts it
emits follow current [`NCATSTranslator/translator-ingests`](https://github.com/NCATSTranslator/translator-ingests)
practice, and each build can emit the Resource Ingest Guide (RIG) metadata those submissions require.

## Getting help

- **Questions and bug reports:** open an issue at
  [github.com/SkyeAv/Tablassert/issues](https://github.com/SkyeAv/Tablassert/issues) (use the
  `question` label for usage questions). Include the command you ran, its output, and your
  `tablassert --version`.
- **Contributing code or docs:** see [CONTRIBUTING.md](CONTRIBUTING.md) for setup, the local quality
  gates, and pull-request expectations. By participating in this project you agree to follow the
  [Code of Conduct](CODE_OF_CONDUCT.md).
- **Reference:** the [CLI Reference](https://skyeav.github.io/Tablassert/cli/) and
  [configuration guides](https://skyeav.github.io/Tablassert/configuration/graph/) document every
  flag and field.

## Quick Start

```bash
pip install "tablassert[cli]"
```

Given a CSV of gene-disease associations with p-values and sample sizes, declare the mapping in a
table config (`table.yaml`):

```yaml
template:
  source:
    kind: text
    local: ./gene-disease.csv
    url: [https://example.com/data.csv]
    row_slice: [1, auto]
    delimiter: ","
  statement:
    subject: { method: column, encoding: A, prioritize: [Gene] }
    predicate: associated_with
    object: { method: column, encoding: B, prioritize: [Disease] }
  provenance: { repo: PMID, publication: "12345678" }
  annotations:
    - { annotation: p_value, method: column, encoding: C }
    - { annotation: study_size, method: column, encoding: D }
```

Wrap it in a graph config (`graph.yaml`) pointing at your fullmap entity-resolution database
and carrying the required `rig:` metadata for the generated Resource Ingest Guide. Build or
download that database once first (`tablassert build-fullmap`, a multi-GB download; see the
[Fullmap guide](https://skyeav.github.io/Tablassert/fullmap/)), which makes `./fullmap` below the
right path:

```yaml
name: MY_KG
version: 1.0.0
tables:
  - ./table.yaml
fullmap: ./fullmap
rig:
  source_info:
    infores_id: infores:my-kg
    terms_of_use_info:
      terms_of_use_url: https://example.org/terms
    data_access_locations:
      - My source downloads - https://example.org/downloads
    source_status: maintained_regular_updates
  ingest_info:
    utility: Gene-disease associations support Translator disease-mechanism queries.
    scope: Gene-disease associations extracted from tabular sources.
  provenance_info:
    contributions:
      - "Author Name - code author, data modeling"
  artifact_base_url: https://example.org/my-kg
  artifact_base_path: ./published/my-kg
```

Build the knowledge graph:

```bash
tablassert build-kg graph.yaml
```

Output is one JSON object per line: nodes with Biolink categories, edges with annotations.

```json
{"id":"HGNC:11998","name":"TP53","category":["biolink:Gene"],"taxon":"NCBITaxon:9606"}
{"id":"MONDO:0008903","name":"lung cancer","category":["biolink:Disease"]}
```

```json
{"id":"2cfea591-0f8f-33af-a7df-03da531d3359","subject":"HGNC:11998","predicate":"biolink:associated_with","object":"MONDO:0008903","p_value":"1.0000e-03","statistical_significance_qualifier":"strongly_significant","has_supporting_studies":{"PMID:12345678":{"id":"PMID:12345678","name":"gene-disease.csv","study_size":450,"has_study_results":[{"id":"row:2"}]}},"publications":["PMID:12345678"]}
```

See the [Tutorial](https://skyeav.github.io/Tablassert/tutorial/) for the full walkthrough.

## Key Features

- **Declarative YAML configuration**: define data transformations without writing code
- **Built-in entity resolution**: map free text to genes, diseases, and chemicals with standard
  CURIEs, taxonomic filtering, and provenance, backed by an embedded redb database
- **Optional quality control**: a four-stage audit (exact -> fuzzy -> abbreviation -> SapBERT embeddings) flags
  low-confidence mappings
- **KGX compliance**: emits NCATS Translator-compatible node/edge NDJSON with Biolink categories
  and predicates
- **Autonomous agent (experimental)**: `tablassert agent` derives, builds, and refines configs for
  whole papers
- **Performance & reproducibility**: lazy Polars pipelines and a deterministic UV-based
  development environment

## Installation

```bash
pip install tablassert
```

Or with uv: `uv tool install "tablassert[cli]"`. The base install provides the Python API; `[cli]`
provides the `tablassert` command. Other extras are opt-in:

| Extra | Adds | Install |
| ----- | ---- | ------- |
| `cli` | `tablassert` command and rich terminal progress | `pip install "tablassert[cli]"` |
| `rt` | CPU-compatible Polars runtime | `pip install "tablassert[rt]"` |
| `aria2` | bundled aria2c downloader, used automatically by `build-fullmap` when installed (Linux/Windows wheels only) | `pip install "tablassert[aria2]"` |
| `qc` | four-stage QC audit (exact -> fuzzy -> abbreviation -> SapBERT embeddings) | `pip install "tablassert[qc]"` |
| `agent` | autonomous agent (smolagents, litellm, article/table context) -- experimental, API may change | `pip install "tablassert[agent]"` |
| `optimize` | GEPA prompt optimization for `agent --optimize` (dspy) -- experimental, API may change | `pip install "tablassert[optimize]"` |
| `distill` | distillation dataset export (`tablassert distill-export`, HF `datasets`) -- experimental, API may change | `pip install "tablassert[distill]"` |
| `log` | loguru-backed file/progress logging (rotation, enqueue) | `pip install "tablassert[log]"` |

Reaching a feature whose extra is missing never produces a bare `ModuleNotFoundError`: the failure
names the missing package and the install command that fixes it. QC is opt-in at build time
(`build-kg --qc`). See the
[Installation guide](https://skyeav.github.io/Tablassert/installation/) for the full matrix,
per-command preflight behavior, and the
[CLI Reference](https://skyeav.github.io/Tablassert/cli/) for every flag.

## Entity Resolution API

```python
from pathlib import Path
from tablassert.lib import resolve_many

results = resolve_many(
    col="gene",
    entities=["TP53", "BRCA1"],
    fullmap=Path("/path/to/fullmap"),
    taxon="9606",
)
# [{"original_gene": "TP53", "gene": "HGNC:11998", "gene_name": "TP53", ...}, ...]
```

Point `resolve_many()` at a fullmap database to resolve any iterable of entity strings to CURIEs,
no LazyFrame setup or NLP preprocessing required. See the
[Batch Resolution API](https://skyeav.github.io/Tablassert/api/lib/) for the full reference.

## Documentation

- **[Installation](https://skyeav.github.io/Tablassert/installation/)**: install methods, extras, and development setup
- **[Tutorial](https://skyeav.github.io/Tablassert/tutorial/)**: step-by-step example with synthetic data
- **[Use Case Gallery](https://skyeav.github.io/Tablassert/examples/)**: real-world configuration patterns
- **[CLI Reference](https://skyeav.github.io/Tablassert/cli/)**: complete command-line flag reference
- **[Fullmap](https://skyeav.github.io/Tablassert/fullmap/)**: building and querying the entity-resolution database
- **[Agent](https://skyeav.github.io/Tablassert/agent/)**: the autonomous agent pipeline
- **[Configuration](https://skyeav.github.io/Tablassert/configuration/graph/)**: graph and table configuration reference
- **[API Reference](https://skyeav.github.io/Tablassert/api/fullmap/)**: core functions documentation
- **[Development](https://skyeav.github.io/Tablassert/development/)**: dev environment setup and contributor workflow
- **[Changelog](https://skyeav.github.io/Tablassert/changelog/)**: release history

## Developing

```bash
uv sync --group dev --extra cli --extra qc --extra log
uv run maturin develop --manifest-path rust/Cargo.toml
make check
```

See **[CONTRIBUTING.md](CONTRIBUTING.md)** for the full development loop, quality gates, and pull
request guidelines.

## Citation

If you use Tablassert, please cite it as described in [CITATION.cff](CITATION.cff). The approach is
described in:

> Skye Lane Goetz, Amy K. Glen, and Gwênlyn Glusman. "MicrobiomeKG: bridging microbiome research
> and host health through knowledge graphs." *Frontiers in Systems Biology* 5 (2025).
> [doi:10.3389/fsysb.2025.1544432](https://doi.org/10.3389/fsysb.2025.1544432)

## License

[Apache License 2.0](LICENSE)

## Contributors

- [Skye Lane Goetz](mailto:sgoetz@isbscience.org), Institute for Systems Biology
- [Gwênlyn Glusman](mailto:gglusman@isbscience.org), Institute for Systems Biology
- Jared C. Roach, Institute for Systems Biology

