Metadata-Version: 2.4
Name: hugpy-wrapper
Version: 0.1.2
Summary: The hugpy wrapper (import name: fitevict): a per-box evict-to-fit front door. An OpenAI /v1 server that resolves a called model to its weights (GGUF via llama.cpp, or transformers), places it on the GPU by evicting what doesn't fit, serves it, and logs every call.
Author-email: putkoff <support@hugpy.ai>
License: Proprietary
Project-URL: Homepage, https://hugpy.ai
Keywords: llama.cpp,inference,evict-to-fit,vram,openai,wrapper,allocator
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Operating System :: POSIX :: Linux
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.9
Description-Content-Type: text/markdown
Provides-Extra: test
Requires-Dist: pytest; extra == "test"
Provides-Extra: db
Requires-Dist: psycopg[binary]>=3; extra == "db"
Provides-Extra: transformers
Requires-Dist: torch==2.14.0; extra == "transformers"
Requires-Dist: transformers==5.17.0; extra == "transformers"
Requires-Dist: accelerate==1.15.0; extra == "transformers"
Requires-Dist: bitsandbytes==0.50.2; extra == "transformers"
Requires-Dist: peft==0.20.0; extra == "transformers"
Requires-Dist: compressed-tensors>=0.15.0; extra == "transformers"
Requires-Dist: safetensors==0.8.0; extra == "transformers"
Requires-Dist: tokenizers==0.23.2; extra == "transformers"
Requires-Dist: sentencepiece==0.2.2; extra == "transformers"
Requires-Dist: tiktoken==0.14.0; extra == "transformers"
Requires-Dist: protobuf==7.36.1; extra == "transformers"
Requires-Dist: huggingface_hub==1.31.0; extra == "transformers"

# hugpy-wrapper (`fitevict`)

    pip install hugpy-wrapper                 # import fitevict; fitevict-serve
    pip install "hugpy-wrapper[transformers]" # the transformers engine (own venv)

hugpy's per-box wrapper: an **evict-to-fit front door over llama.cpp** (and transformers). Point any OpenAI
client's base URL at it; it turns one address into a self-managing model server.

A `/v1` call carries only a model **name** — never a location. llama has no idea
where a gguf lives. `fitevict` is the layer that does what the call can't: for
each request it resolves the name → gguf path, runs an **evict-to-fit** plan
against what's resident and what's being called, places the model into llama
(launch/evict), forwards the call, splits the model's `<think>` out of the
answer, and records the call — raw stamps + engine geometry — to a call log that
every metric derives from.

The package depends on nothing else in hugpy. The decision core is pure stdlib; hardware
facts come from `nvidia-smi`; model/gguf facts from the files on disk.

## Layout

```
fitevict/
  types.py          frozen data contract (DeviceBudget, Resident, LoadRequest, FitPlan, ...)
  evict.py          plan_eviction / sort_key  — the victim selector (pure)
  plan.py           plan_fit                   — staged decision + quant ladder (pure)
  flex.py           ctx-band compress + layers-that-fit offload (pure)
  host.py           EngineHost protocol + drive() loop (measure→plan→evict→load)
  front_door.py     the WRAPPER: persistent OpenAI /v1 server (owns the address)
  adapters/
    llama_cpp.py    concrete EngineHost: measure GPU/RAM, price GGUFs, launch/evict llama-server
    db.py           call_log writer (Postgres inference_engine.call_log)
  lifted/           measure/act helpers lifted clean from hugpy (imports stripped):
                    gguf_inspect, gguf_need, hardware, spill_reserve, model_resolve,
                    native_resolve, supervisor/procutil, timings, no_think, app_dirs, ...
  run_engine.py     a juggling harness: discover models, fire a sequence that exceeds the card
```

## Install

```
pip install .            # core + wrapper (stdlib only)
pip install .[db]        # + psycopg for direct call_log DSN writes (else shells to psql)
pip install .[test]      # + pytest
```

## Run

```
fitevict-serve --host 127.0.0.1 --port 8080     # the front door (owns /v1)
fitevict-run                                     # the juggle harness on this box
```

Then point any OpenAI client at `http://127.0.0.1:8080/v1`:

```
curl -s http://127.0.0.1:8080/v1/chat/completions -H 'Content-Type: application/json' \
  -d '{"model":"<name>","messages":[{"role":"user","content":"hi"}],"max_tokens":128}'
```

The model is named, not located; the wrapper finds it, fits it, serves it, logs it.

## Test

```
pytest            # pure unit tests + a call_log replay + the VL-flood fixture
```
