Metadata-Version: 2.5
Name: uubed
Version: 1.0.6
Summary: High-performance, position-safe embedding encoding
Project-URL: Homepage, https://github.com/twardoch/uubed
Project-URL: Documentation, https://github.com/twardoch/uubed/blob/main/README.md
Project-URL: Repository, https://github.com/twardoch/uubed.git
Project-URL: Bug Tracker, https://github.com/twardoch/uubed/issues
Author-email: Adam Twardoch <adam+github@twardoch.com>
License-Expression: MIT
License-File: LICENSE
Keywords: SIMD,embedding,encoding,rust,search,vector
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Rust
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Requires-Python: >=3.10
Requires-Dist: click>=8.0
Requires-Dist: defusedxml>=0.7.1
Requires-Dist: numpy>=1.20
Requires-Dist: rich>=10.0
Requires-Dist: toml>=0.10.0
Requires-Dist: typing-extensions>=4.15.0
Provides-Extra: all
Requires-Dist: chromadb>=0.4.0; extra == 'all'
Requires-Dist: cupy-cuda11x>=10.0; extra == 'all'
Requires-Dist: langchain>=0.1.0; extra == 'all'
Requires-Dist: pinecone-client>=2.0; extra == 'all'
Requires-Dist: pytest-cov>=4.0; extra == 'all'
Requires-Dist: pytest>=7.0; extra == 'all'
Requires-Dist: qdrant-client>=1.0; extra == 'all'
Requires-Dist: weaviate-client>=3.0; extra == 'all'
Provides-Extra: gpu
Requires-Dist: cupy-cuda11x>=10.0; extra == 'gpu'
Provides-Extra: inference
Requires-Dist: einops>=0.8.1; extra == 'inference'
Requires-Dist: httpx>=0.28; extra == 'inference'
Requires-Dist: llama-cpp-python>=0.3.35; extra == 'inference'
Requires-Dist: sentence-transformers<6,>=5.2; extra == 'inference'
Requires-Dist: transformers<5,>=4.57.6; extra == 'inference'
Provides-Extra: langchain
Requires-Dist: langchain>=0.1.0; extra == 'langchain'
Provides-Extra: mlx
Requires-Dist: mlx-embeddings==0.0.5; (platform_machine == 'arm64' and sys_platform == 'darwin') and extra == 'mlx'
Requires-Dist: mlx-vlm==0.3.9; (platform_machine == 'arm64' and sys_platform == 'darwin') and extra == 'mlx'
Provides-Extra: onnx
Requires-Dist: sentence-transformers[onnx]<6,>=5.2; extra == 'onnx'
Provides-Extra: test
Requires-Dist: pytest-cov>=4.0; extra == 'test'
Requires-Dist: pytest>=7.0; extra == 'test'
Provides-Extra: tm
Requires-Dist: uubed-rs>=1.0.12; extra == 'tm'
Provides-Extra: vectordb
Requires-Dist: chromadb>=0.4.0; extra == 'vectordb'
Requires-Dist: pinecone-client>=2.0; extra == 'vectordb'
Requires-Dist: qdrant-client>=1.0; extra == 'vectordb'
Requires-Dist: weaviate-client>=3.0; extra == 'vectordb'
Description-Content-Type: text/markdown

---
this_file: README.md
---
# Uubed Python

Shared model inference, compact embedding records and English-source translation
memory for book alignment and localization.

## Translation memory

Install this checkout together with its native `uubed-rs` wheel. The `[tm]` extra
requires that distribution; add `[inference]` for local embedding models. See the
[workspace setup](../../README.md) for editable installs before publication.

```python
from pathlib import Path
from uubed.embeddings import Embedder
from uubed.memory import TranslationMemory, build_memory

engine = Embedder("jina", dimensions=256, device="cpu")
try:
    build_memory([Path("translations.tmx")], Path("localization.sqlite"), engine)
    with TranslationMemory("localization.sqlite", embedder=engine) as memory:
        exact = memory.exact("Bold", "pl")
        references = memory.lookup("Make the font bold", "pl", top_k=5)
finally:
    engine.close()
```

The importer streams TMX, finds English by language tag, and embeds each distinct
source once. SQLite stores all target variants and path/unit provenance, including
conflicting translations. Exact lookup preserves case and whitespace and loads no
model. Semantic lookup filters by target language and returns bounded cosine-ranked
pairs. Model/weights/task/prefix/truncation/dimension mismatches are rejected.
The index uses ordinary SQLite and signed int8 vectors, without a vector extension.
At 256 dimensions, each vector payload occupies 256 bytes; texts and metadata add
storage. Build into a new filename to change inputs or model settings.

```bash
uubed tm build translations.tmx --output localization.sqlite --model jina --dimensions 256
uubed tm lookup localization.sqlite 'Bold' --language pl --exact-only
uubed tm lookup localization.sqlite 'Make the font bold' --language pl
```

For a GGUF-backed index, pass `--backend llama.cpp --model-path /path/to/model.gguf`
at build time and the matching `--model-path` at semantic lookup time.
The index can move independently of the weight file.

## Book embeddings

[EMBEDDINGS.md](EMBEDDINGS.md) describes the shared `Embedder`, verified model
profiles, resident runtimes, Matryoshka truncation and UB1 int8/int4/binary records.
Vexy Paraltext uses these APIs for its book alignment workflow and TMX output.
TM indexing uses the same inference and signed int8 quantization code.

## Byte codecs

The existing `encode`/`decode` APIs retain `eq64`, `shq64`, `t8q64`, `zoq64`, and
`mq64`. They encode bytes and are separate from model-aware signed UB1 vectors.
The optional native module is named `uubed_native`; Python fallbacks remain
available for byte encoding. A textual lossless encoding expands bytes; it is not
an embedding compression algorithm.

```python
from uubed import encode, decode
encoded = encode(bytes(range(256)), method="eq64")
decoded = decode(encoded, method="eq64")
```

## Development

`../../test.sh` builds the actual native wheel, installs it into the test
environment, and verifies this package and both consumers. `uvx hatch test -py 3.12`
runs this package's suite when its native TM dependency is installed.
Build distributions with `uv build`. MIT license.

## Releases and local data

`./publish.sh --dry-run` verifies the next release without pushing or uploading.
`./publish.sh` commits, tags and publishes it. See [RELEASING.md](RELEASING.md)
for credentials, same-tag retries, dependency order and private-data exclusions.
