Metadata-Version: 2.4
Name: hdw-dedup-engine
Version: 0.1.0
Summary: Reusable duplicate-image detection engine using corroborated perceptual hashing.
Author: Hans De Weme
License-Expression: MIT
Project-URL: Homepage, https://code2trade.dev/
Project-URL: Repository, https://github.com/hansdeweme/HdWDedupEngine
Project-URL: Issues, https://github.com/hansdeweme/HdWDedupEngine/issues
Keywords: duplicate-images,perceptual-hashing,image-deduplication,dhash,phash,whash
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Topic :: Multimedia :: Graphics
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: numpy>=1.26
Requires-Dist: Pillow>=10.0
Provides-Extra: heic
Requires-Dist: pillow-heif>=0.16; extra == "heic"
Provides-Extra: ssim
Requires-Dist: scikit-image>=0.22; extra == "ssim"
Provides-Extra: trash
Requires-Dist: send2trash>=1.8; extra == "trash"
Provides-Extra: full
Requires-Dist: pillow-heif>=0.16; extra == "full"
Requires-Dist: scikit-image>=0.22; extra == "full"
Requires-Dist: send2trash>=1.8; extra == "full"
Provides-Extra: dev
Requires-Dist: pytest>=8; extra == "dev"
Requires-Dist: pytest-cov>=5; extra == "dev"
Requires-Dist: build>=1.2; extra == "dev"
Requires-Dist: twine>=5; extra == "dev"
Dynamic: license-file

# HdWDedupEngine

HdWDedupEngine is a reusable Python library for detecting duplicate and near-duplicate
images. It provides the shared engine used by DedupTool and ChronoName, without
depending on their GUIs or source trees.

- **Distribution:** `hdw-dedup-engine`
- **Python import:** `hdw_dedup_engine`
- **Current version:** `0.1.0`
- **Requires:** Python 3.10 or later
- **License:** MIT

## Features

- Perceptual matching using dHash, with pHash and wHash corroboration.
- Optional exact-file hashing and SSIM checks.
- Duplicate clustering and keeper selection.
- CSV and HTML reports.
- Separate quarantine/trash actions with dry-run support.
- Optional incremental indexing, progress callbacks and cancellation.

Scanning produces a plan; moving files is a separate action controlled by the
calling application. Applications also own their settings, report and cache paths.

## Local installation

From this repository:

```powershell
py -3 -m pip install -e .
```

Or install directly from another project:

```powershell
py -3 -m pip install -e D:\Coding\HdWDedupEngine
```

The package is not assumed to be available on PyPI. No DedupTool checkout or
global `PYTHONPATH` setting is needed.

Core dependencies are NumPy and Pillow. Optional extras are:

| Extra | Capability |
| --- | --- |
| `heic` | HEIC decoding through pillow-heif |
| `ssim` | SSIM through scikit-image |
| `trash` | Operating-system trash through send2trash |
| `full` | All three optional capabilities |
| `dev` | Tests and distribution build tools |

For example:

```powershell
py -3 -m pip install -e ".[heic,trash,dev]"
```

## Basic usage

This example analyzes a folder and writes a CSV report without moving images:

```python
import os
from pathlib import Path

from hdw_dedup_engine import (
    DedupConfig,
    DedupRunOptions,
    load_settings,
    plan_duplicates,
    write_csv,
)

root = Path(r"D:\Photos\TestCollection").resolve()
reports = root / "Duplicate Reports"
settings = load_settings()
# Make the application-owned report location explicit and exclude it from scans.
settings.setdefault("reports", {})["base_dir"] = os.path.normcase(str(reports))
settings.setdefault("scan", {}).setdefault("exclude_roots", []).append(str(reports))

result = plan_duplicates(DedupRunOptions(
    roots=[root],
    config=DedupConfig(),
    settings=settings,
    log=print,
))

print(f"Scanned {result.files_scanned} images; {result.cluster_count} clusters")
print(f"Keep: {len(result.keep_paths)}; duplicates: {len(result.drop_paths)}")
write_csv(result.raw_summary, str(reports / "duplicates.csv"), [str(root)])
```

Review the plan before using `execute_moves`. Its default is `dry_run=True`;
actual moves require explicitly setting `dry_run=False`. HTML reports use
`write_html` with the same raw summary. Quarantine destinations preserve relative
paths and avoid overwriting existing files.

## Public API

Import application-facing APIs from the package root:

- `DedupConfig`, `DedupRunOptions`, `DedupPlanResult`
- `plan_duplicates`, `create_engine`, `DedupEngine`
- `evaluate_pair`, `MatchEvidence`
- `load_settings`, `write_csv`, `write_html`, `execute_moves`

Scan progress callbacks receive `(done, total, phase)`. Action progress callbacks
receive a single percentage from 0 to 100. Cancellation callbacks return a boolean.
Advanced consumers can supply an `IndexDB` from `hdw_dedup_engine.index_db`; its
lifetime and storage location remain the caller's responsibility.

## Tests and distribution builds

```powershell
py -3 -m pip install -e ".[dev]"
py -3 -m pytest -v
py -3 -m build
```

Build artifacts are written to `dist/`. The regression suite currently contains
79 tests; optional-dependency tests may be skipped when their extras are absent.

## Integration notes

Version 0.1.0 retains the existing duplicate-matching and keeper policies.
`load_settings` chooses the working-directory settings file in source mode and
the executable-directory settings file when frozen; its supplied path argument
is currently ignored. Consumers should supply explicit absolute report/cache
paths and exclude their output directories. Windows report paths should be
normalized with `os.path.normcase`.

For PyInstaller applications, collect `hdw_dedup_engine` and the optional
dependencies the application uses. GUI resources and other external tools are
packaged by the consuming application.
