Metadata-Version: 2.5
Name: otb-kbo-opendata-pyclient
Version: 2.0.0
Summary: Library and kbo-opendata command to discover and download KBO Open Data files from the Belgian KBO SFTP server.
Project-URL: Homepage, https://github.com/openthebox/kbo-opendata-pyclient
Project-URL: Repository, https://github.com/openthebox/kbo-opendata-pyclient
Project-URL: Issues, https://github.com/openthebox/kbo-opendata-pyclient/issues
Project-URL: Changelog, https://github.com/openthebox/kbo-opendata-pyclient/blob/main/CHANGELOG.md
Author-email: openthebox <vincent.kox@openthebox.be>
License-Expression: MIT
License-File: LICENSE
Keywords: belgium,cbe,cli,kbo,opendata,sftp
Classifier: Development Status :: 5 - Production/Stable
Classifier: Environment :: Console
Classifier: Intended Audience :: Developers
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Utilities
Classifier: Typing :: Typed
Requires-Python: >=3.10
Requires-Dist: boto3>=1.28
Requires-Dist: paramiko>=3.0
Description-Content-Type: text/markdown

# otb-kbo-opendata-pyclient

[![CI](https://github.com/openthebox/kbo-opendata-pyclient/actions/workflows/ci.yml/badge.svg?branch=main)](https://github.com/openthebox/kbo-opendata-pyclient/actions/workflows/ci.yml)
[![Release](https://github.com/openthebox/kbo-opendata-pyclient/actions/workflows/release.yml/badge.svg)](https://github.com/openthebox/kbo-opendata-pyclient/actions/workflows/release.yml)
[![Coverage](https://img.shields.io/endpoint?url=https%3A%2F%2Fgist.githubusercontent.com%2Fv-kox%2F878ef07a4cc1e80bfc028ff100dbcb5b%2Fraw%2Fcoverage.json)](https://github.com/openthebox/kbo-opendata-pyclient/actions/workflows/ci.yml)
[![PyPI](https://img.shields.io/pypi/v/otb-kbo-opendata-pyclient.svg)](https://pypi.org/project/otb-kbo-opendata-pyclient/)
[![Python versions](https://img.shields.io/pypi/pyversions/otb-kbo-opendata-pyclient.svg)](https://pypi.org/project/otb-kbo-opendata-pyclient/)
[![Licence](https://img.shields.io/pypi/l/otb-kbo-opendata-pyclient.svg)](https://github.com/openthebox/kbo-opendata-pyclient/blob/main/LICENSE)

Python client for the SFTP server of the Belgian Crossroads Bank for Enterprises (KBO/BCE), which
publishes daily KBO Open Data ZIP files. Use it as a library or through the `kbo-opendata` command.

## Features

- Discover the latest publication, or look one up by index or by date
- List everything on the server, filter by index or date range, and spot skipped indices
- Download by filename, index, date, or "the latest", choosing update files, full files, or both
- Sync a destination with the server, fetching only the files it does not already hold
- Write to a local directory or straight to S3, streaming rather than buffering whole files
- Host-key verification against `known_hosts` by default
- Fully typed, with a `py.typed` marker

## Installation

```bash
pip install otb-kbo-opendata-pyclient
```

## Getting access

KBO Open Data is free but not anonymous. Two steps are needed before this client can connect:

1. Register on the [KBO Open Data portal](https://kbopub.economie.fgov.be/kbo-open-data) and accept
   the [licence](https://economie.fgov.be/sites/default/files/Files/Entreprises/BCE/Licence-BCE-Open-Data-Conditions-d-utilisation.pdf).
   This alone gives you manual downloads through the website.
2. Request SFTP access separately, in advance, by emailing
   [kbo-bce-webservice@economie.fgov.be](mailto:kbo-bce-webservice@economie.fgov.be). Portal
   registration does not grant it.

The credentials you receive are what this package uses. There is no test or sandbox server.

## Credentials

The client needs the username and password issued by KBO. The library accepts them directly; the
CLI reads them from the environment.

| Variable | Purpose |
| --- | --- |
| `KBO_OPENDATA_USERNAME` | SFTP account name, which also names the remote directory |
| `KBO_OPENDATA_PASSWORD` | SFTP password |

## About the published files

KBO snapshots its database daily and publishes two ZIP files per snapshot:

- a **full** file — every active registered entity and establishment unit at the moment of the snapshot
- an **update** file — the differences between the last full file and the one before it

**Files are kept for 31 days only.** Anything older is gone from the server, so `list_all()`,
`missing_indices()` and date lookups only ever see roughly the last month. Download the full file
first, then keep up with either update files or a periodic full file.

The `XXXX` in a filename is the `ExtractNumber`, incremented by one per publication. It is not a
date: KBO can skip a calendar day, which is why `missing_indices()` reports gaps in the sequence
rather than missing dates.

### What is inside a ZIP

This package downloads and stores the ZIP files; it does not extract or parse them. Each archive
holds CSV files — `meta.csv`, `code.csv`, `enterprise.csv`, `establishment.csv`, `denomination.csv`,
`address.csv`, `contact.csv`, `activity.csv` and `branch.csv` — joined on the enterprise number,
establishment number or branch id. The CSV conventions are a comma delimiter, double-quoted text,
a full stop as decimal point, and `dd-mm-yyyy` dates.

`meta.csv` carries `SnapshotDate`, `ExtractTimestamp`, `ExtractType` (`full` or `update`),
`ExtractNumber` and the format `Version`, which is the reliable way to confirm what an archive
actually contains.

An update archive splits each table into a `_delete` and an `_insert` file. Applying one means
deleting every row for the listed entity numbers, then inserting the rows from the `_insert` file —
the insert file repeats all current rows for a changed entity, not only the changed ones. The files
carry no history: only the current state of active entities.

### Documentation

- [Cookbook](https://economie.fgov.be/sites/default/files/Files/Entreprises/BCE/Cookbook-BCE-Open-Data.pdf)
  — file structure, CSV field descriptions and the update procedure
- Data catalogue, the reusable data fields — no English version exists, only
  [Dutch](https://economie.fgov.be/sites/default/files/Files/Entreprises/KBO/Gegevenscatalogus-hebruikbare-gegevens-KBO-Open-Data.pdf)
  and [French](https://economie.fgov.be/sites/default/files/Files/Entreprises/BCE/Catalogue-des-donnees-reutilisables-BCE-opendata.pdf)
- [KBO Open Data page](https://economie.fgov.be/en/themes/enterprises/crossroads-bank-enterprises/services-everyone/public-data-available-reuse/cbe-open-data)

## Library usage

```python
import datetime as dt

from kbo_opendata import KboOpenDataClient, LocalDestination, S3Destination

with KboOpenDataClient(username="...", password="...") as client:
    latest = client.latest()
    print(latest.update_filename, latest.full_filename)

    pair = client.get_by_index(423)  # None when the index was skipped
    pair = client.get_by_date(dt.date(2026, 8, 7))

    client.download_latest(LocalDestination("./downloads"))
    client.download_index(423, S3Destination("my-bucket", "kbo/"), kinds=["full"])
```

`KboOpenDataClient.from_env()` builds the same client from the environment variables above.

### Queries

| Method | Returns |
| --- | --- |
| `latest()` | The pair with the highest index, or `None` |
| `latest_n(count)` | The most recent pairs, newest last |
| `get_by_index(index)` | The pair for that index, or `None` |
| `get_by_date(date)` | The pair for that date, or `None` |
| `list_all()` | Every pair, oldest first |
| `list_by_index_range(start, end)` | Pairs within inclusive index bounds |
| `list_by_date_range(start, end)` | Pairs within inclusive date bounds |
| `missing_indices()` | Indices skipped between the lowest and highest present |
| `exists(filename)` | Whether the server holds that file |
| `stat(filename)` | Size and modification time |
| `catalogue(refresh=False)` | The whole listing as a `Catalogue`, fetched once and reused |
| `refresh()` | Fetch the listing again and return it |

A `KboFilePair` carries `update_filename` and `full_filename`, either of which is `None` when the
server holds only one of the two. Both are bare filenames, without a path.

The listing is fetched on first use and cached, so repeated queries cost nothing; call `refresh()`
to pick up files published since. Use the client as a context manager, as above, or call `close()`
to release the connection and drop the cached listing. The `store` property exposes the underlying
transport, which is what the test suite substitutes.

### Downloads

| Method | Downloads |
| --- | --- |
| `download(filenames, destination)` | The named files |
| `download_index(index, destination)` | The files for one index |
| `download_date(date, destination)` | The files for one date |
| `download_latest(destination)` | The most recent files |

Every download method accepts `overwrite` (defaults to `False`, so existing files are skipped) and
`dry_run`. The index, date and latest variants also accept `kinds`, which defaults to update and
full; `download` takes explicit filenames, so it has no `kinds`. They all return a `DownloadResult`
whose `written`, `skipped`, `missing` and `planned` tuples say what happened to each file.

The index and date variants raise `RemoteFileNotFoundError` when nothing matches; `download` reports
unknown names as `missing` instead, so a multi-file request always reports on every name.

### S3

```python
S3Destination("my-bucket", "kbo/", extra_args={"ServerSideEncryption": "AES256"})
```

A boto3 client is created from the ambient AWS configuration on first use. Pass `client=` to supply
one built from your own session, and `extra_args` to forward parameters to the upload.

## Command line usage

```bash
export KBO_OPENDATA_USERNAME=...
export KBO_OPENDATA_PASSWORD=...

kbo-opendata --version
kbo-opendata show-latest
kbo-opendata check-index 0423
kbo-opendata check-date 2026-08-07
kbo-opendata list-files --from-index 400 --to-index 425
kbo-opendata list-files --from-date 2026-08-01 --to-date 2026-08-07 --latest 10
kbo-opendata list-gaps
kbo-opendata check-credentials

kbo-opendata download-latest --dest ./downloads
kbo-opendata download-indexes 421 423 --dest ./downloads --kind full
kbo-opendata download-dates 2026-08-03 2026-08-07 --s3-bucket my-bucket --s3-prefix kbo/
kbo-opendata download KboOpenData_0423_2026_08_07_Full.zip --dest ./downloads

kbo-opendata sync --dest ./downloads
```

`sync` downloads the files on the server that are not yet at the destination. Presence is decided
on the filename alone, so a file whose content changed on the server is not fetched again unless
`--overwrite` is given. It honours `--kind`, so it collects both kinds unless told otherwise, and
it ignores remote entries that are not KBO data files. Its text output lists what it transfers and
then counts what was already there, rather than naming every file again:

```text
KboOpenData_0425_2026_08_07_Full.zip: written -> ./downloads/KboOpenData_0425_2026_08_07_Full.zip
848 already present, 1 written (412.7 MB)
```

Dates are accepted as `2026-08-07`, `2026_08_07` or `20260807`, on every supported Python version.

Add `--json` for machine-readable output, `--overwrite` to replace existing files, `--dry-run` to
see what would be transferred, `-v`/`-vv`/`-q` to adjust logging, and `--version` to print the
version and exit.

### JSON output

`--json` replaces the text output with a single JSON document. The download commands report every
file they considered:

```json
{
  "outcomes": [
    {
      "filename": "KboOpenData_0421_2026_08_03_Full.zip",
      "status": "written",
      "location": "./downloads/KboOpenData_0421_2026_08_03_Full.zip",
      "size": 431495168
    }
  ],
  "written": 1,
  "skipped": 0,
  "missing": 0,
  "planned": 0,
  "total_bytes": 431495168,
  "unresolved": ["index 0423"]
}
```

`status` is `written`, `skipped`, `missing` or `planned`, and `location` and `size` are `null`
where they do not apply. `unresolved` names the requests that matched nothing on the server; only
`download-indexes` and `download-dates` fill it, and it holds labels such as `index 0423` or
`2026-08-05` rather than filenames. Unlike the text output, `--json` always reports every file,
including the ones `sync` would otherwise only count.

The other commands are shaped as follows.

| Command | Payload |
| --- | --- |
| `show-latest`, `check-index`, `check-date` | `{"index", "date", "update", "full"}`, or `null` when nothing matched |
| `list-files` | An array of those same objects |
| `list-gaps` | `{"missing_indices": [423, 427]}` |
| `check-credentials` | `{"files", "pairs", "latest_index"}` |

In those payloads `date` is an ISO date, while `update` and `full` are bare filenames, either of
which is `null` when the server holds only one of the two.

### Exit codes

| Code | Meaning |
| --- | --- |
| `0` | Success |
| `1` | Nothing matched the request, or a batch matched only in part: one or more requested files or values were absent from the server |
| `2` | Usage error |
| `3` | Configuration, authentication or connection failure, or any other client error |
| `4` | The destination could not be written to |

## Host keys

The client verifies the server's host key against `~/.ssh/known_hosts` and refuses to connect to an
unknown host. Add the server once with `ssh-keyscan`, or pass `accept_unknown_host_key=True`
(`--accept-unknown-host-key` on the CLI) to trust it on first use.

## Logging

The package logs through the standard `logging` module under the `kbo_opendata` logger and installs
no handlers of its own. The CLI configures a stderr handler for its own process.

## Development

```bash
uv sync --group dev
uv run ruff check . && uv run ruff format --check .
uv run mypy
uv run coverage run -m pytest && uv run coverage report
```

The test suite is fully offline: the SFTP transport and the download destinations sit behind typed
protocols, and the tests drive real in-memory implementations of them rather than mocks.

## Reusing the data

The data is covered by the KBO Open Data licence, which you accept at registration — separate from
this package's MIT licence, which covers only the code. One restriction is worth stating plainly:
**personal data from these files may not be reused for direct marketing purposes.** See the
[licence](https://economie.fgov.be/sites/default/files/Files/Entreprises/BCE/Licence-BCE-Open-Data-Conditions-d-utilisation.pdf)
and the [CBE privacy statement](https://economie.fgov.be/en/themes/enterprises/crossroads-bank-enterprises/crossroads-bank-enterprises-0).

## Licence

MIT
