Metadata-Version: 2.5
Name: signal-dataset
Version: 0.3.0
Summary: An open, profile-free format and Python library for large-scale multidimensional signal datasets.
Project-URL: Documentation, https://superpose-labs.github.io/signal-dataset/
Project-URL: Issues, https://github.com/superpose-labs/signal-dataset/issues
Project-URL: Repository, https://github.com/superpose-labs/signal-dataset
Author: Superpose Labs
License-Expression: Apache-2.0
License-File: LICENSE
Requires-Python: <3.14,>=3.11
Requires-Dist: array-record<0.9,>=0.8.3
Requires-Dist: numpy<2.5,>=1.26
Requires-Dist: safetensors<0.9,>=0.5
Provides-Extra: gcs
Requires-Dist: google-cloud-storage<4,>=3; extra == 'gcs'
Provides-Extra: s3
Requires-Dist: boto3<2,>=1.35.69; extra == 's3'
Description-Content-Type: text/markdown

# Signal Dataset

Immutable, indexed storage for multidimensional signal records on local filesystems, GCS, and S3.
Numerical fields use SafeTensors; ArrayRecord provides random access.

> **Status:** experimental 0.x software. The Python API may make documented breaking changes
> between minor releases. Persisted-format compatibility is versioned separately.

```python
import numpy as np
import signal_dataset as sds

record = sds.Record(
    id="capture-0042",
    fields={
        "iq": sds.Field(
            np.zeros((4, 4096), dtype=np.complex64),
            axes=(sds.Axis("channel", 4), sds.Axis("time", 4096)),
        )
    },
    metadata={"sample_rate_hz": 20_000_000},
)

shard = sds.write_shard([record], "captures.sds", work_id="worker-000")
dataset = sds.publish(
    "captures.sds",
    [shard],
    dataset_id="captures",
    snapshot_id="run-001",
)

assert dataset[0].id == "capture-0042"
assert dataset.record_metadata[0]["iq"].shape == (4, 4096)

for descriptor in dataset.iter_record_metadata():
    print(descriptor.id)

for full_record in dataset.iter_records():
    assert full_record["iq"].data.shape == (4, 4096)
```

`dataset.record_metadata[index]` reads one aligned metadata entry without fetching signal tensors.
`dataset.iter_record_metadata()` streams all metadata in logical order with bounded shard-store
requests and keeps memory bounded to `StorageOptions.records_per_read_batch` records. The shipped
ArrayRecord store serves each request with one underlying reader lifetime; custom stores control
their own `read_many()` implementation.
`dataset.iter_records()` provides the same ordered, shard-batched traversal for full records,
including tensor payloads and normal record validation.
Readers follow generation-pinned manifests and never list directories or GCS prefixes.

See the [quickstart](docs/quickstart.md), [annotation example](docs/how-to/annotations.md), and
[distributed-writing guide](docs/how-to/distributed-writing.md).

## Install

```bash
uv add signal-dataset
uv add 'signal-dataset[gcs]'  # GCS transport
uv add 'signal-dataset[s3]'   # S3 transport
```

Equivalent pip commands are `pip install signal-dataset`,
`pip install 'signal-dataset[gcs]'`, and `pip install 'signal-dataset[s3]'`.

Python 3.11–3.13 on Linux and macOS are supported. Windows is not currently supported because the
local atomic-publication implementation uses POSIX filesystem primitives.

## Design

- Immutable, root-last publication: readers observe no dataset or a complete snapshot.
- Lazy ordinal random access without directory or bucket-prefix discovery.
- Tensor-free aligned metadata reads.
- Storage and shard-container extension contracts.
- Profile-free records: no modality, training framework, or split policy is embedded in the core.

The project does not provide mutable datasets, distributed scheduling, authorization, retention,
or batching policy. See the [architecture](docs/architecture.md), [format compatibility
contract](docs/concepts/compatibility.md), and [roadmap](ROADMAP.md).

Signal Dataset is use-case agnostic. It has no training-framework, modality, or split-policy
dependency.

## Releasing

Releases publish to PyPI through GitHub Actions using [trusted
publishing](https://docs.pypi.org/trusted-publishers/), so no API token is stored anywhere. The
`Release` workflow is triggered by pushing a `v*` tag and refuses to publish unless the tag, the
package version, and the changelog agree.

To cut a release from `main`:

1. Bump `version` in `pyproject.toml`.
2. Add a dated `## <version> — YYYY-MM-DD` section to `CHANGELOG.md`. The workflow rejects a
   heading still marked `— Unreleased`.
3. Update the version assertion in `test_library_release_does_not_change_persisted_layout_version`
   (`tests/test_foundation.py`). It pins the library version next to `LAYOUT_VERSION` so that
   changing one makes you confirm the other deliberately.
4. Run `uv lock` so the lockfile records the new version.
5. Merge to `main`, then tag the merge commit and push it:

   ```bash
   git tag -a v0.3.0 -m "signal-dataset 0.3.0"
   git push origin v0.3.0
   ```

The workflow then re-runs every gate — tag/version match, the changelog heading, the full test
suite with the coverage threshold, `ruff`, `mypy`, `mkdocs build --strict`, `uv build`, an
installed-wheel version check, and the base/`[gcs]`/`[s3]`/sdist install journeys — before
uploading. If any gate fails, nothing reaches PyPI.

Rehearse the tag and changelog gates locally before pushing:

```bash
uv version --short                       # must equal the tag without its leading v
grep -F "## $(uv version --short)" CHANGELOG.md
```

### Trusted publisher configuration

Publishing requires a trusted publisher registered on the PyPI **project**, at
`https://pypi.org/manage/project/<name>/settings/publishing/`. A *pending* publisher added under
account settings only applies to projects that do not exist yet and will not match an existing one.

| Field | Value |
| --- | --- |
| Owner | the GitHub organization, e.g. `superpose-labs` |
| Repository name | the repository alone, e.g. `signal-dataset` |
| Workflow name | the filename only: `release.yml` |
| Environment name | `pypi`, matching `environment:` in the workflow |

A mismatch surfaces as `invalid-publisher: valid token, but no corresponding publisher` in the
publish step. That message prints the exact claims GitHub sent, which is what the fields above
must match.

## Community

Read [CONTRIBUTING.md](CONTRIBUTING.md) before proposing changes. Use GitHub Issues for reproducible
bugs and design discussion, and follow [SECURITY.md](SECURITY.md) for private vulnerability reports.
Participation is governed by the [Code of Conduct](CODE_OF_CONDUCT.md).
