Metadata-Version: 2.4
Name: fit_changedetector
Version: 0.1.0a1
Summary: Compare geo-data, report on diffs
Author-email: Simon Norris <snorris@hillcrestgeo.ca>
Project-URL: Homepage, https://github.com/bcgov/fit_changedetector
Project-URL: Issues, https://github.com/bcgov/fit_changedetector
Classifier: Development Status :: 1 - Planning
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Topic :: Scientific/Engineering :: GIS
Classifier: Operating System :: OS Independent
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: click
Requires-Dist: cligj
Requires-Dist: geopandas>=1.0.0
Requires-Dist: pyarrow
Provides-Extra: test
Requires-Dist: pytest; extra == "test"
Requires-Dist: build; extra == "test"
Requires-Dist: pre-commit; extra == "test"
Dynamic: license-file

# FIT Change Detector

[![Lifecycle:Maturing](https://img.shields.io/badge/Lifecycle-Maturing-007EC6)](https://github.com/bcgov/repomountie/blob/master/doc/lifecycle-badges.md)

Compare two sets of geo-data and report on the differences.

## Installation

Install with pip:

    pip install fit_changedetector

An ArcGIS Pro script tool is also provided (`arcgis.py`).
Because an ArcGIS managed conda environment is unlikely to be 100% compatible with this module's dependencies, installation of this module to a virtual environment is recommended:

In a Windows Command Prompt (with no active conda environment):

    python -m venv .venv
    .venv\Scripts\activate.bat
    pip install fit_changedetector

Set the `FIT_CHANGEDETECTOR_VENV_PYTHON` environment variable to the path of `python.exe` in your virtual environment (e.g. via `setx FIT_CHANGEDETECTOR_VENV_PYTHON "C:\path\to\.venv\Scripts\python.exe"`, or through Windows' System Properties > Environment Variables), then drop `arcgis.py` into your ArcGIS toolbox. If you don't have permission to set an environment variable, instead create a `venv_python.txt` file next to `arcgis.py` in the toolbox folder, containing just the path to `python.exe`. To avoid conflict with the system Python, the script tool passes the arguments provided in the ArcGIS tool to the change detector CLI - which is run in a subprocess using the virtual environment's Python.


## Usage

#### Python module

The primary function of interest is `gdf_diff()`:

    import geopandas
    import fit_changedetector as fcd

    # read the data
    df_a = geopandas.read_file(in_file_a, layer=layer_a)
    df_b = geopandas.read_file(in_file_b, layer=layer_b)

    # compare the two dataframes
    diff = fcd.gdf_diff(
        df_a,
        df_b,
        <primary_key>,
        fields=<fields_to_compare>,
        precision=<precision>,
        suffix_a="a",
        suffix_b="b",
    )

`gdf_diff` returns a dictionary having the keys noted below. Dictionary values are geopandas GeoDataFrames holding the corresponding records.

Dictionary keys:

| key | description |
|-----|-------------|
| `NEW` | additions |
| `DELETED` | deleted records |
| `UNCHANGED` | unchanged records |
| `MODIFIED_BOTH` | records where attribute columns and geometries have changed |
| `MODIFIED_ATTR` | records where attribute columns have changed but geometries have not changed |
| `MODIFIED_GEOM` | records where geometries have changed but attribute columns have not |
| `DUPLICATES` | records dropped due to a duplicated primary key (only populated if `allow_duplicates=True`) |

Schemas for records contained in `NEW`, `DELETED`, `UNCHANGED` are as per the source data. `DUPLICATES` records retain their source schema too, plus a `_fcd_source_` column (`suffix_a`/`suffix_b`) identifying which source each record was dropped from.
Schemas for records contained in the `MODIFIED` keys include only columns where a change has occurred.
For example, these are some "modified attributes" records, with "_a" suffix for values from the primary dataset, and "_b" suffix for values from the secondary dataset:

    >>> diff["MODIFIED_ATTR"]
      id       park_name_a                park_name_b parkclasscode_a parkclasscode_b
    0  3  Mars Street Park        Jupiter Street Park             NaN             NaN
    1  6      Mayfair Blue              Mayfair Green              BL             GRN
    2  7                    Quadra Heights Playground             NaN             NaN
    3  9               NaN                        NaN              RP             PND


#### CLI

<!-- BEGIN CLI HELP (auto-generated by scripts/update_cli_docs.py - do not edit by hand) -->
    $ changedetector --help
    Usage: changedetector [OPTIONS] COMMAND [ARGS]...

    Options:
      --version  Show the version and exit.
      --help     Show this message and exit.

    Commands:
      add-hash-key  Read input data, compute hash, write to new file
      diff          Compare two datasets, printing a JSON summary to stdout
      diff2gdb      Compare two datasets, writing results to .gdb

    $ changedetector add-hash-key --help
    Usage: changedetector add-hash-key [OPTIONS] IN_FILE OUT_FILE

      Read input data, compute hash, write to new file

    Options:
      --in-layer TEXT           Name of layer to add hashed primary key
      -nln, --out-layer TEXT    Output layer name
      -hk, --hash-key TEXT      Name of new column containing hashed data
      -d, --drop-null-geometry  Drop records with null geometry
      -hf, --hash-fields TEXT   Comma separated list of fields to include in the
                                hash (not including geometry)
      -p, --precision FLOAT     Coordinate precision for geometry hash and
                                comparison. Default=0.01
      --crs TEXT                Coordinate reference system to use when hashing
                                geometries (eg EPSG:3005)
      -v, --verbose             Increase verbosity.
      -q, --quiet               Decrease verbosity.
      --help                    Show this message and exit.

    $ changedetector diff --help
    Usage: changedetector diff [OPTIONS] IN_FILE_A IN_FILE_B

      Compare two datasets, printing a JSON summary to stdout

      Same comparison as `diff2gdb`, but for when spatial output isn't needed -
      prints a JSON summary instead of writing a .gdb: record counts per
      NEW/DELETED/UNCHANGED/MODIFIED_* category, plus the primary key value(s)
      present in each category (use --count to omit the key lists and print just the
      counts).

      IN_FILE_A may be "-" to read GeoJSON from stdin instead of a file.

    Options:
      --layer-a TEXT             Name of layer to use within in_file_a (not valid if
                                 reading from stdin/parquet)
      --layer-b TEXT             Name of layer to use within in_file_b (not valid if
                                 reading from parquet)
      -f, --fields TEXT          Comma separated list of fields to compare (do not
                                 include primary key)
      -if, --ignore-fields TEXT  Comma separated list of fields to ignore
      -pk, --primary-key TEXT    Comma separated list of primary key column(s),
                                 common to both datasets
      -hk, --hash-key TEXT       Name of new column to add as hash key
      -hf, --hash-fields TEXT    Comma separated list of fields to include in the
                                 hash (in addition to geometry)
      -p, --precision FLOAT      Coordinate precision for geometry hash and
                                 comparison. Default=0.01
      -a, --suffix-a TEXT        Suffix to append to column names from data source A
                                 when comparing attributes
      -b, --suffix-b TEXT        Suffix to append to column names from data source B
                                 when comparing attributes
      -d, --drop-null-geometry   Drop records with null geometry
      --crs TEXT                 Coordinate reference system to use when hashing
                                 geometries (eg EPSG:3005)
      -c, --count                Print only record counts, omitting the primary key
                                 values in each category
      --allow-duplicates         Do not fail on a duplicated primary key - instead,
                                 drop all but the first occurrence of each
                                 duplicated key from the source it was found in, and
                                 include a DUPLICATES category in the output
      -v, --verbose              Increase verbosity.
      -q, --quiet                Decrease verbosity.
      --help                     Show this message and exit.

    $ changedetector diff2gdb --help
    Usage: changedetector diff2gdb [OPTIONS] IN_FILE_A IN_FILE_B

      Compare two datasets, writing results to .gdb

      IN_FILE_A may be "-" to read GeoJSON from stdin instead of a file.

    Options:
      --layer-a TEXT             Name of layer to use within in_file_a (not valid if
                                 reading from stdin/parquet)
      --layer-b TEXT             Name of layer to use within in_file_b (not valid if
                                 reading from parquet)
      -f, --fields TEXT          Comma separated list of fields to compare (do not
                                 include primary key)
      -if, --ignore-fields TEXT  Comma separated list of fields to ignore
      -o, --out-file PATH        Path to output file, defaults to
                                 ./changedetector_YYYYMMDD_HHMM.gdb
      -pk, --primary-key TEXT    Comma separated list of primary key column(s),
                                 common to both datasets
      -hk, --hash-key TEXT       Name of new column to add as hash key
      -hf, --hash-fields TEXT    Comma separated list of fields to include in the
                                 hash (in addition to geometry)
      -p, --precision FLOAT      Coordinate precision for geometry hash and
                                 comparison. Default=0.01
      -a, --suffix-a TEXT        Suffix to append to column names from data source A
                                 when comparing attributes
      -b, --suffix-b TEXT        Suffix to append to column names from data source B
                                 when comparing attributes
      -d, --drop-null-geometry   Drop records with null geometry
      -i, --dump-inputs          Dump input layers (with new hash key) to output
                                 .gdb
      --crs TEXT                 Coordinate reference system to use when hashing
                                 geometries (eg EPSG:3005)
      --allow-duplicates         Do not fail on a duplicated primary key - instead,
                                 drop all but the first occurrence of each
                                 duplicated key from the source it was found in, and
                                 write the dropped records to a DUPLICATES layer
      -v, --verbose              Increase verbosity.
      -q, --quiet                Decrease verbosity.
      --help                     Show this message and exit.
<!-- END CLI HELP -->

##### Examples

Compare the test datasets using their known primary key:

    $ changedetector diff2gdb -v \
        tests/data/parks_a.geojson \
        tests/data/parks_b.geojson \
        -pk id

Compare the test datasets, using a hash of geometry and the column `park_name` as synthetic primary key, written to `new_hash_column`:

    $ changedetector diff2gdb -v \
        tests/data/parks_a.geojson \
        tests/data/parks_b.geojson \
        -hf park_name \
        -hk new_hash_column

`IN_FILE_A` may be `-` to read GeoJSON from stdin instead of a file, e.g. to compare a database export against a file on disk without writing the export to disk first:

    $ ogr2ogr -f GeoJSON /vsistdout/ PG:"dbname=mydb" -sql "SELECT * FROM my_table" | \
        changedetector diff2gdb -v - tests/data/parks_b.geojson -pk id

`diff` prints record counts per category plus the primary key value(s) present in each by default; add `--count`/`-c` to print just the counts:

    $ changedetector diff tests/data/parks_a.geojson tests/data/parks_b.geojson -pk id
    {"NEW": 1, "DELETED": 1, "UNCHANGED": 1, "MODIFIED_BOTH": 1, "MODIFIED_ATTR": 4, "MODIFIED_GEOM": 1, "keys": {"NEW": ["8"], "DELETED": ["2"], "UNCHANGED": ["1"], "MODIFIED_BOTH": ["5"], "MODIFIED_ATTR": ["3", "6", "7", "9"], "MODIFIED_GEOM": ["4"]}}

    $ changedetector diff tests/data/parks_a.geojson tests/data/parks_b.geojson -pk id --count
    {"NEW": 1, "DELETED": 1, "UNCHANGED": 1, "MODIFIED_BOTH": 1, "MODIFIED_ATTR": 4, "MODIFIED_GEOM": 1}

##### Usage notes

Layer options are not valid for stdin and parquet sources - streams and parquet files have no concept of multiple layers.

A *partitioned* parquet dataset (a directory of many `.parquet` files) is not supported as a single source - `diff`/`diff2gdb` load an entire source into memory regardless of format, so there's no benefit to teaching them to read a partition directory as one dataset. Instead, run the command once per file, e.g. for two partitioned datasets with matching partition filenames:

    $ for part in dataset_a/*.parquet; do
        name=$(basename "$part" .parquet)
        changedetector diff2gdb -v \
            "$part" \
            "dataset_b/${name}.parquet" \
            -pk id \
            -o "output_${name}.gdb"
      done

Neither `diff` nor `diff2gdb` have a bounding box / spatial filtering option, and won't - applying a bbox filter independently to each source risks reporting spurious `NEW`/`DELETED` records for anything that moved across the boundary between the two snapshots being compared, rather than genuinely appearing/disappearing from the source. Filter with another tool first, e.g.:

    $ ogr2ogr -f GeoJSON /vsistdout/ dataset_a.gpkg -spat 1150000 470000 1200000 500000 | \
        changedetector diff2gdb -v - dataset_b.gpkg -pk id


#### ArcGIS

The script tool calls the above documented CLI. Documentation of the parameters is also provided within the ArcGIS interface.

## Subtleties to geometry change detection

Prior to comparing geometries, the tool will:

- normalize geometries (vertex order/starting point or ring winding direction)
- promote mixed single/multipart types to multipart (when a source contains both variants — a uniformly single-part source compared against a uniformly multi-part source will still raise a type mismatch)
- apply coordinate precision tolerance (`-p`/`--precision`, default 0.01)

On the other hand, a new vertex in an otherwise unchanged geometry will be considered as MODIFIED_GEOM.
See geopandas [geom_equals_exact](https://geopandas.org/en/stable/docs/reference/api/geopandas.GeoSeries.geom_equals_exact.html) for details.

Curved geometry types (`CIRCULARSTRING`, `COMPOUNDCURVE`, `CURVEPOLYGON`, etc, as found in some `.gdb` sources) are not supported - only `POINT`, `LINESTRING`, `POLYGON` and their `MULTI*` equivalents are ([#66](https://github.com/bcgov/FIT_changedetector/issues/66)). GDAL segments curves into their linear approximation on read, so a curved source may appear to work, but the precision of that approximation is not controlled by `diff`/`diff2gdb` and results involving curved input should not be relied on.



## Development and testing

Uses [uv](https://docs.astral.sh/uv/) for dependency management:

    $ git clone git@github.com:bcgov/FIT_changedetector.git
    $ cd FIT_changedetector
    $ uv sync --extra test
    $ uv run pytest
