Metadata-Version: 2.4
Name: premixdb
Version: 0.1.0
Summary: Data mixtures for LLM pretraining that you can reproduce and trace
Requires-Python: >=3.12
Description-Content-Type: text/markdown
Requires-Dist: protobuf<7,>=6.33.5
Requires-Dist: blake3>=1.0.10
Requires-Dist: packaging>=24
Requires-Dist: ipython<10,>=9.17.1
Requires-Dist: tokenizers<0.24,>=0.23.1
Requires-Dist: datatrove==0.3.0; sys_platform == "darwin" and platform_machine == "x86_64"
Requires-Dist: datatrove==0.10.1; sys_platform != "darwin" or platform_machine != "x86_64"
Requires-Dist: numpy<2,>=1.26; sys_platform == "darwin" and platform_machine == "x86_64"
Requires-Dist: fasteners>=0.19
Requires-Dist: fasttext-wheel==0.9.2; sys_platform == "darwin" and platform_machine == "x86_64"
Requires-Dist: fasttext-numpy2-wheel>=0.9.2; sys_platform != "darwin" or platform_machine != "x86_64"
Requires-Dist: regex>=2025.10.22
Requires-Dist: tldextract<6,>=5.3
Requires-Dist: spacy<4,>=3.8
Requires-Dist: fsspec[http]>=2023.12.0
Requires-Dist: marin-dupekit==0.1.1
Requires-Dist: pyarrow<24,>=18
Requires-Dist: transformers<6,>=5.10
Requires-Dist: sentence-transformers<6,>=5.7
Requires-Dist: torch<2.3,>=2.2; sys_platform == "darwin" and platform_machine == "x86_64"
Requires-Dist: torch<3,>=2.5; sys_platform != "darwin" or platform_machine != "x86_64"
Requires-Dist: einops<1,>=0.8
Requires-Dist: datasets<6,>=5.0.1
Provides-Extra: huggingface
Requires-Dist: datasets<6,>=5.0.1; extra == "huggingface"

# premixdb

premixdb is a declarative language for building reproducible data mixtures to improve llm pretraining.

```bash
uv add premixdb
```

## Capture, filter, mix, train

```python
import premixdb as p
from torch.utils.data import DataLoader

with p.PremixDB(storage=".premixdb") as db:
    c4 = db.corpus(
        "c4",
        p.HuggingFaceSource("datablations/c4-filter-small"),
        limit=8,
    )
    oscar = db.corpus(
        "oscar",
        p.HuggingFaceSource("datablations/oscar-filter-small"),
        limit=8,
    )
    query = c4.union(oscar).query(
        steps=[
            p.where(p.text.characters >= 200),
            p.dedupe(),
        ]
    )
    print(query.preview())

    mixtures = query.mix(tokens=256, sequence_length=64)
    print(mixtures.weights)
    print(mixtures.profile())

    dataset = mixtures[0]
    print(dataset.preview())
    batch = next(iter(DataLoader(dataset.torch(), batch_size=1)))
```

A mix creates three candidate datasets by default. Repeat a recipe to reuse its
results. Keep `.premixdb` to reopen your corpora and datasets.

## Use the fields you need

```python
with p.PremixDB(storage=".premixdb") as db:
    query = db.corpus("c4").query(
        steps=[
            p.where(p.language.en >= 0.75),
            p.where(p.quality.educational_value >= 1.0),
        ]
    )
    print(query.profile())
    print(query.preview())
```

Fields are computed on demand and cached. Model fields download their assets on
first use. The default dataset tokenizer is bundled GPT-2; use your model's
tokenizer for training.

## Explore

```bash
uvx --python 3.12 premixdb shell
```

The shell opens with `db`, `p`, and an offline Shakespeare `demo` corpus.
From a checkout, use `make shell` (`uvx --from . premixdb shell`).
Use `make install` for an editable local CLI install.

- [Runnable examples](examples/README.md)
- [Curation](docs/curation.md) · [Fields](docs/enrichment.md) · [Training](docs/training.md)
- [Storage](docs/persistence.md) · [Internals](docs/architecture.md) · [Development](docs/development.md)
