Metadata-Version: 2.1
Name: custc
Version: 0.1.0
Summary: CUDA-accelerated Syndrome-Trellis Codes: batched GPU ±1 embedding for steganography research
Keywords: steganography,STC,syndrome-trellis-codes,CUDA,GPU
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: C++
Classifier: Environment :: GPU :: NVIDIA CUDA
Classifier: Intended Audience :: Science/Research
Classifier: Topic :: Security
Requires-Python: >=3.9
Requires-Dist: numpy>=1.20
Provides-Extra: examples
Requires-Dist: matplotlib>=3.5; extra == "examples"
Provides-Extra: jpeg
Requires-Dist: matplotlib>=3.5; extra == "jpeg"
Requires-Dist: jpeglib>=1.0; extra == "jpeg"
Requires-Dist: conseal>=2024.1; extra == "jpeg"
Description-Content-Type: text/markdown

# cuSTC

CUDA-accelerated Syndrome-Trellis Codes (STC) for steganography.

STCs embed a message into a cover object while minimizing an additive distortion function, using a Viterbi decoder over a syndrome trellis. This project implements the Viterbi core (forward Add-Compare-Select pass + traceback) on the GPU, running many independent trellises in parallel — one CUDA block per trellis. Output is bit-identical to the reference CPU implementation, so messages embedded on the GPU are extracted by the standard CPU extractor and vice versa.

The original CPU implementation from the Digital Data Embedding Laboratory (Binghamton University) is included, with the Boost dependency removed in favor of the C++11 `<random>` header, as done in [pySTC](https://github.com/daniellerch/pySTC).

## Features

- Batched GPU embedding: `B` independent trellises in one call, one CUDA block each — with per-trellis `(n, m)` shapes (`stc_embed_cuda_batch_v`), so heterogeneous datasets and per-image payload splits batch together.
- Batched ±1 multi-layer embedding (`stc_pm1_pls_embed_cuda_batch`): the full `stc_pm1_pls_embed` construction — permutation, payload split by cost entropy, two embedding layers, infeasibility retries — restructured per-layer across the batch. Pixels + per-pixel ±1 costs + payload in, stego pixels + distortion out; extraction is the unchanged reference `stc_ml_extract()`.
- Automatic strategy selection per call: a fast 1-pass traceback when the path fits in VRAM, a checkpoint/recompute strategy (`O(sqrt(n))` memory) when it does not. Both produce identical output.
- Automatic wave sizing from free VRAM and the kernel's measured occupancy; multi-wave batches are double-buffered on two CUDA streams so transfers overlap compute.
- Distortion-only mode (`stego = nullptr`): skips traceback and path memory, useful for payload search.
- Wet-pixel support via large costs; infeasible embeddings are reported (`distortion = -1`), never silently corrupted.
- No external dependencies beyond the CUDA toolkit.

## Requirements

- NVIDIA GPU and CUDA toolkit (tested with CUDA 12.4 on an RTX 2060).
- `g++` with SSE2 (for the CPU reference code).

## Install (Python)

```sh
pip install .   # needs the CUDA toolkit (nvcc) on PATH
```

```python
import numpy as np, custc

# covers: list of int arrays (2-D images fine); costs: (n, 3) float per image
# with the cost of the {-1, 0, +1} change per pixel; messages: 0/1 bit arrays
res = custc.embed_pm1_batch(covers, costs, messages, h=10)   # GPU, one call
assert not res.img_err.any()                                 # per-image status
msgs = custc.extract_batch(res.stegos, res.num_msg_bits, h=10)  # CPU, exact
```

`embed_pm1_batches()` consumes a generator of mini-batches for datasets that
don't fit host RAM; `embed_binary_batch()` exposes the raw binary Viterbi
core. Embedding releases the GIL; device errors raise, per-image
infeasibility is reported in `img_err`.

**Demos** — both run the whole pipeline end-to-end (pull a few BOSSbase images by
HTTP range request, no 1.6 GB download; compute costs; embed the batch on the GPU;
extract; verify exact; plot the change map):

- [`example/bossbase_hill_demo.ipynb`](example/bossbase_hill_demo.ipynb) — spatial
  domain, HILL costs.
- [`example/jpeg_juniward_demo.ipynb`](example/jpeg_juniward_demo.ipynb) — JPEG
  domain: compress to real JPEGs, embed in the quantized DCT coefficients with
  J-UNIWARD costs (via [`conseal`](https://pypi.org/project/conseal/), derived from
  the DDE reference), write stego JPEGs and recover the message from the re-read
  files. Needs `pip install jpeglib conseal`.

## Build (C/C++ library)

```sh
cmake -S . -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build -j        # libcustc.so + libcustc.a, headers in ml_stc_src
```

The default is a fat binary (SASS for sm_75/80/86/89/90 + PTX); for faster
development builds target just your GPU with `-DCMAKE_CUDA_ARCHITECTURES=75`.

The Makefile builds the examples and benchmarks:

```sh
make            # GPU + CPU examples   (set ARCH, default sm_75: make ARCH=sm_86)
make benchmarks # main_compare, main_benchmark_cuda + the verification harnesses
make test       # run the full verification suite (needs a GPU, see Testing)
```

## Usage (C/C++)

**±1 embedding into pixel data** — the entry point research pipelines use — is one call, declared in `ml_stc_src/stc_ml_cuda.h`:

```c
// B images, per-image lengths n_arr[i]; covers/stegos are pixel values,
// costs is 3 floats per pixel {cost of -1, 0, +1}, messages are bit arrays.
// Per-image failures are reported in img_err without failing the batch.
int stc_pm1_pls_embed_cuda_batch(int B, const int* n_arr, const int* covers,
                                 const float* costs, const int* msg_arr,
                                 const u8* messages, int h, float wet_cost,
                                 uint max_trials, int* stegos, float* out_dist,
                                 uint* num_msg_bits, uint* trials, int* img_err);
```

The binary Viterbi core is exposed in `ml_stc_src/stc_cuda.cuh`:

```c
// single trellis: returns distortion, or -1 if the message is not embeddable
double stc_embed_cuda(const u8* cover, int n, const u8* msg, int m,
                      const float* cost, u8* stego, int h);

// B independent trellises with per-trellis (n, m); data concatenated.
// Returns 0, or nonzero on device failure; out_dist[b] = -1 marks an
// infeasible syndrome. stc_embed_cuda_batch is a uniform-shape wrapper.
int stc_embed_cuda_batch_v(int B, const u8* cover, const int* n_arr,
                           const u8* msg, const int* m_arr, const float* cost,
                           u8* stego, double* out_dist, int h);
```

Binary inputs are cover bits (typically an LSB plane), per-element flip costs (use a large value for wet pixels), the message bits, and the constraint height `h` (2^h trellis states; larger `h` embeds closer to the theoretical bound but costs more time). Extraction is unchanged: use the reference `stc_extract()` / `stc_ml_extract()`.

For throughput, prefer one batched call over a loop of single embeds — the GPU is filled by running many trellises at once. Batch sizing, VRAM budgeting, and stream overlap are handled internally.

For a loop over many mini-batches, use a **plan** (cuFFT-style) so the device-property/occupancy queries, H-structure build, and buffer allocation happen once instead of per call:

```c
// declare the shapes you expect and a batch-size hint (both optional);
// device = -1 uses the current device, h is fixed for the plan's lifetime
stc_plan plan;
stc_plan_create(&plan, n_shapes, m_shapes, num_shapes, h, max_B, /*device=*/-1);
for (...)  // reuses device buffers, streams and the pinned pool across calls
    stc_plan_embed(plan, B, cover, n_arr, msg, m_arr, cost, stego, out_dist);
stc_plan_destroy(plan);
```

`stc_embed_cuda_batch_v` is a create-embed-destroy one-shot over this API. A plan is bound to one device and one `h`; multi-GPU is one plan per device. All device work runs on the plan's own non-blocking streams — no device-wide synchronization — so it composes with a PyTorch/multi-stream pipeline.

## Examples

| Example | Shows |
|---|---|
| `example1_gpu` (`ml_stc_src/example1_gpu_stc.cpp`) | Single-cover GPU embedding with wet pixels, verified with the CPU extractor |
| `example2_gpu` (`ml_stc_src/example2_gpu_stc.cpp`) | Batched embedding, distortion-only mode, throughput measurement |
| `example1_ml_stc`, `example2_ml_stc` | Original CPU multi-layer ±1 / ±2 constructions (Filler's reference) |
| `example/bossbase_hill_demo.ipynb` | Python, spatial-domain: BOSSbase + HILL costs, GPU embed, extract, change map |
| `example/jpeg_juniward_demo.ipynb` | Python, JPEG-domain: DCT coefficients + J-UNIWARD costs, stego-JPEG roundtrip |

Run the C/C++ examples with `./example.sh`; open the notebooks in Jupyter.

## Testing

Bit-identical-to-reference is the project's contract, and it is a runnable test, not a claim:

```sh
make test       # needs a GPU; exit code 0 iff every case passes
```

This builds and runs four harnesses (all in `benchmark/`):

| Harness | Checks |
|---|---|
| `main_verify` | The correctness bar itself: identical total distortion vs the CPU reference (exact, not within tolerance) and 0-error extraction, swept over `h ∈ {10..13}` × `alpha ∈ {0.25, 0.33, 0.5}`, both traceback strategies (every cell runs under auto dispatch and forced checkpoint, plus a B=1 huge-n case where auto dispatch must pick checkpoint itself), odd `n`, and wet costs in both conventions (`INFINITY` and `1e13`) |
| `main_verify_v` | Mixed-shape batches (`stc_embed_cuda_batch_v`): duplicated shapes, multi-wave pipeline, infeasible trellises mid-batch, distortion-only mode, batched output bit-identical to B=1 calls, `h` up to 14 (gmem kernel) |
| `main_verify_ml` | The ±1 multi-layer path (`stc_pm1_pls_embed_cuda_batch`) vs the CPU `stc_pm1_pls_embed`, image by image: identical payload splits, trial counts, distortion and stego pixels; exact extraction; per-image failures |
| `main_verify_plan` | The plan API vs the one-shot path: byte-identical across buffer growth, new shapes, waves, two live plans, device selection |

CI (`.github/workflows/ci.yml`) compiles everything — Makefile targets and the CMake fat binary — on every push inside NVIDIA's CUDA container (no GPU needed); the GPU `make test` job targets a self-hosted runner and is enabled by setting the repository variable `CUSTC_GPU_RUNNER=true` once a runner with labels `[self-hosted, gpu]` is registered.

## Conventions & reproducibility

**Wet costs.** A wet element (one that must not be modified) is marked by its cost, and the two common conventions behave differently — pick deliberately:

- `INFINITY` (the reference's `F_INF`): flipping is impossible. If the message cannot be embedded without touching wet elements, the embedder reports the trellis as infeasible (`distortion = -1`) instead of silently corrupting one. Internally a syndrome is infeasible when its optimal metric reaches `0.5 * FLT_MAX`, which only infinite costs produce.
- A large finite cost such as `1e13` (HStego's value): flipping is merely astronomically expensive. The embedder will flip a wet element rather than fail, and the returned distortion then includes that cost; infeasibility is never reported. Avoid finite values near `FLT_MAX`, which alias the infeasibility guard.

Both conventions are supported and exercised by `make test`. For the ±1 path's `wet_cost` parameter (the forbidden ±2 change), `~1e13` is the HStego-compatible choice.

**Determinism.** For identical inputs the stego output is bit-identical to the CPU reference: both sides accumulate the same f32 costs in the same trellis order (the reference's SSE core computes in f32 too), and both break metric ties toward the flip predecessor. Output is independent of batch composition, wave splitting, traceback strategy, plan reuse, and GPU model — the kernels use no atomics or order-varying reductions — so results reproduce across machines. (In the ±1 construction the host-side payload split runs the reference's own float code, so splits, permutation seeds and retry decisions are bit-identical as well; its binary layers run f32 against the reference's f64 driver, so an exact cost tie could in principle resolve differently — none has been observed across the test matrix.)

**Permutation seeds (±1 path).** As in the reference, each layer's cover permutation is `randperm(n, seed)` with the seed equal to that layer's final `num_msg_bits` value, generated by `std::mt19937` (Boost-free, matching pySTC). A failed trial decrements `num_msg_bits`, which both drops one message bit and selects the next permutation. Extraction re-derives the permutations from `num_msg_bits` alone — that pair is the only side information needed.

## Benchmarks

Everything benchmark-related lives in `benchmark/`. `benchmark/main_compare` embeds identical inputs on GPU and CPU and checks distortion equality and extraction. `./benchmark/run_compare_sweep.sh` runs a full sweep. Results on an RTX 2060 (6 GB), payload 0.5 bpp, 100k-element chunks:

| Cover size | h | CPU (ms) | GPU (ms) | Speedup |
|---|---|---|---|---|
| 1M | 10 | 1433 | 81 | 17.8x |
| 4M | 10 | 5678 | 152 | 37.3x |
| 1M | 11 | 2814 | 136 | 20.8x |
| 1M | 12 | 5968 | 244 | 24.5x |
| 400k | 13 | 4679 | 461 | 10.2x |

Distortion matches the CPU reference exactly and extraction succeeds in every configuration. See `benchmark/BENCHMARK_GPU_vs_CPU.md` for details and `GPU_ARCHITECTURE.md` for the design.

For the ±1 multi-layer construction (correctness covered by the [Testing](#testing) suite): on the RTX 2060, embedding 512×512 images at 0.4 bpp, h=10 runs at ~26 img/s end-to-end vs ~1.1 img/s for the CPU loop (**~24x**, B=96).

## Roadmap

Work to make the GPU batching usable in real steganography research (batched ±1 multi-layer embedding, per-trellis shapes, Python bindings, error reporting, test suite) is tracked in [ROADMAP.md](ROADMAP.md).

## Acknowledgments
Part of the C/C++ code used by HStego comes from the [Digital Data Embedding Laboratory](http://dde.binghamton.edu/download/).
We sincerely appreciate their work in developing and sharing this technology.

These methods are described in the following papers:

- [Minimizing Embedding Impact in Steganography using Trellis-Coded Quantization](https://doi.org/10.1117/12.838002) by Tomas Filler, Jan Judas and Jessica Fridrich.

- [Minimizing Additive Distortion Functions With Non-Binary Embedding Operation in Steganography](https://doi.org/10.1109/WIFS.2010.5711444) by Tomas Filler and Jessica Fridrich.

- [Minimizing Additive Distortion in Steganography using Syndrome-Trellis Codes](https://doi.org/10.1109/TIFS.2011.2134094) by Tomas Filler, Jan Judas and Jessica Fridrich.

## License
This project is licensed under the MIT License. See the [LICENSE](/LICENSE.txt) file for more details.
