Metadata-Version: 2.5
Name: bitnarrow
Version: 0.1.0
Summary: Open weight model surgery lab: edit, measure, gate, and ship LLM transformations
Project-URL: Homepage, https://github.com/amareshhebbar/bitnarrow
Project-URL: Repository, https://github.com/amareshhebbar/bitnarrow
Project-URL: Issues, https://github.com/amareshhebbar/bitnarrow/issues
Project-URL: Models, https://huggingface.co/AmareshHebbar
Author-email: Amaresh Hebbar <hebbar.gvamaresh@gmail.com>
License: Apache-2.0
License-File: LICENSE
Keywords: abliteration,evaluation,huggingface,interpretability,llm,model-surgery,pruning,transformers
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.10
Requires-Dist: huggingface-hub>=0.24
Requires-Dist: pyyaml>=6.0
Provides-Extra: all
Requires-Dist: accelerate>=0.30; extra == 'all'
Requires-Dist: bitsandbytes>=0.43; extra == 'all'
Requires-Dist: datasets>=2.19; extra == 'all'
Requires-Dist: gradio>=4.0; extra == 'all'
Requires-Dist: lm-eval>=0.4.3; extra == 'all'
Requires-Dist: safetensors>=0.4; extra == 'all'
Requires-Dist: torch>=2.2; extra == 'all'
Requires-Dist: transformers>=4.44; extra == 'all'
Provides-Extra: dev
Requires-Dist: accelerate>=0.30; extra == 'dev'
Requires-Dist: build>=1.2; extra == 'dev'
Requires-Dist: pytest>=8.0; extra == 'dev'
Requires-Dist: ruff>=0.5; extra == 'dev'
Requires-Dist: safetensors>=0.4; extra == 'dev'
Requires-Dist: torch>=2.2; extra == 'dev'
Requires-Dist: transformers>=4.44; extra == 'dev'
Requires-Dist: twine>=5.0; extra == 'dev'
Provides-Extra: eval
Requires-Dist: datasets>=2.19; extra == 'eval'
Requires-Dist: lm-eval>=0.4.3; extra == 'eval'
Provides-Extra: merge
Requires-Dist: mergekit>=0.0.5; extra == 'merge'
Provides-Extra: quant
Requires-Dist: bitsandbytes>=0.43; extra == 'quant'
Provides-Extra: space
Requires-Dist: gradio>=4.0; extra == 'space'
Provides-Extra: torch
Requires-Dist: accelerate>=0.30; extra == 'torch'
Requires-Dist: safetensors>=0.4; extra == 'torch'
Requires-Dist: torch>=2.2; extra == 'torch'
Requires-Dist: transformers>=4.44; extra == 'torch'
Description-Content-Type: text/markdown

```
████  ███ █████ █   █  ███  ████  ████   ███  █   █
█   █  █    █   ██  █ █   █ █   █ █   █ █   █ █   █
████   █    █   █ █ █ █████ ████  ████  █   █ █ █ █
█   █  █    █   █  ██ █   █ █  █  █  █  █   █ ██ ██
████  ███   █   █   █ █   █ █   █ █   █  ███  █   █
```

# bitnarrow

Open weight model surgery lab. Edit a model, measure it, gate it, ship it.

bitnarrow takes an open weight LLM, applies a structural transformation from a declarative recipe, scores the base and the result with the same evaluation harness, refuses to ship anything that fails the release gate, computes a weight level diff, and writes a HuggingFace model card with every number in it. One command from recipe to publishable artifact.

```bash
pip install "bitnarrow[all]"
bitnarrow run recipes/narrow/qwen2.5-1.5b-narrow.yaml
bitnarrow publish runs/qwen2.5-1.5b-narrow
```

## Why

Converting a base model for a product (smaller, faster, less over cautious, more specialized) breaks things silently. bitnarrow makes every conversion measured and reproducible:

* **Recipes**: one YAML file fully describes base model, method, data, eval, gate, and output repo. Fingerprinted and copied into every run.
* **Eval spine**: refusal (over refusal on safe prompts, refusal retained on harmful prompts), capability (any lm-evaluation-harness task, plus built in perplexity), and systems (params, weight size, tokens per second, peak VRAM). Base results are cached by eval fingerprint.
* **Release gate**: harmful prompt refusal must stay at or above the base model, capability may not drop beyond a tolerance, and each recipe can add its own requirements. `publish` refuses gate failures.
* **Diff**: tensor level report of what changed (relative change, cosine, fraction changed, optional spectral rank check), per layer and per module.
* **Ship small**: weight edits ship as patches holding only the modified tensors, which hot load onto the stock base model, including 4bit loading.

## Methods

| method | what it does | artifact |
|---|---|---|
| `narrow` | Training free directional edit. Extracts a behavior direction from contrastive prompt sets with Winsorized activation means, discovers the layer and band, projects the direction out of writer (`o_proj`, `down_proj`) and reader (`gate_proj`, `up_proj`) weights, and sweeps edit strength, keeping the strongest edit that passes the gate. | patch |
| `prune` | Depth pruning. Ranks contiguous blocks (or single layers) by angular distance between their input and output residual streams on a calibration set and removes the least influential. | full model |
| `merge` | SLERP, TIES, DARE and friends through mergekit, scored against the base. | full model |
| `none` | Scores the base only, for baselines. | none |

The shipped `narrow` recipe targets over refusal: the direction is built from safe prompts that models commonly refuse (OR-Bench hard) against ordinary instructions, and the gate requires over refusal to drop while harmful prompt refusal holds.

## Install

```bash
pip install bitnarrow                 # recipes, gate, cards, hub tooling
pip install "bitnarrow[torch]"        # model loading, surgery, patches
pip install "bitnarrow[torch,eval]"   # plus datasets and lm-evaluation-harness
pip install "bitnarrow[all]"          # plus bitsandbytes and gradio
```

The shipped recipes use XSTest and HarmBench, which are gated on the Hub. Accept their terms on the dataset pages, then:

```bash
hf auth login
```

## Quickstart

Scaffold a recipe for any base model:

```bash
bitnarrow init narrow llama-3.2-3b-narrow --base meta-llama/Llama-3.2-3B-Instruct --license llama3.2
```

Run it:

```bash
bitnarrow run recipes/narrow/llama-3.2-3b-narrow.yaml --strict
```

A run directory looks like this:

```
runs/llama-3.2-3b-narrow/
  recipe.yaml              exact recipe used
  manifest.json            method, env, gate status, chosen strength, sweep, diff summary
  gate.json                every gate check with values
  results/base.json        base scores (schema versioned)
  results/artifact.json    artifact scores
  results/sweep_*.json     one file per swept strength
  diff.json, diff.md       weight diff
  artifact/                patch.safetensors, patch.json, direction.safetensors (or a full model)
  README.md                model card
```

Publish (refuses if the gate failed):

```bash
bitnarrow publish runs/llama-3.2-3b-narrow
```

## Python API

```python
import bitnarrow

model, tokenizer = bitnarrow.load("AmareshHebbar/Qwen2.5-1.5B-Narrow")
model, tokenizer = bitnarrow.load("AmareshHebbar/Qwen2.5-1.5B-Narrow", load_in_4bit=True)
```

`load` reads `patch.json`, loads the recorded base model, and applies the patch. With `load_in_4bit=True` the patched modules stay in full precision and everything else is quantized to NF4.

## CLI

| command | purpose |
|---|---|
| `bitnarrow init METHOD NAME --base MODEL` | scaffold a recipe |
| `bitnarrow run RECIPE [--strict] [--publish]` | surgery, eval, gate, diff, card |
| `bitnarrow eval MODEL [--patch P] [--recipe R]` | score any model or base plus patch |
| `bitnarrow diff BASE OTHER [--spectral]` | weight diff, works on hub ids, folders, or patches |
| `bitnarrow gate BASE.json ARTIFACT.json [--recipe R]` | apply the gate to two result files |
| `bitnarrow card RUN_DIR` | rerender the model card |
| `bitnarrow publish RUN_DIR` | push a gated run to the Hub |
| `bitnarrow export REF --materialize DIR [--gguf Q4_K_M,Q8_0]` | merge a patch into a full checkpoint, convert to GGUF |
| `bitnarrow env` | environment fingerprint |

GGUF export needs a llama.cpp checkout: set `LLAMA_CPP_DIR` or pass `--llama-cpp`.

## Recipe reference

```yaml
name: qwen2.5-1.5b-narrow
method: narrow
seed: 0
base:
  model: Qwen/Qwen2.5-1.5B-Instruct
  dtype: bfloat16
params:
  positive: {dataset: bench-llm/or-bench, config: or-bench-hard-1k, split: train, column: prompt, limit: 128}
  negative: {dataset: tatsu-lab/alpaca, split: train, column: instruction, limit: 128}
  winsor_quantile: 0.95
  layer: auto
  band: all
  strength: [0.25, 0.5, 0.75, 1.0]
  targets: [o_proj, down_proj, gate_proj, up_proj]
eval:
  suites: [refusal, capability, systems]
  refusal:
    over_refusal: {dataset: walledai/XSTest, split: test, column: prompt, filter: {label: safe}}
    harmful: {dataset: walledai/HarmBench, config: standard, split: train, column: prompt}
  capability:
    tasks: [arc_easy, hellaswag, gsm8k]
    limit: 250
gate:
  harmful_refusal_tolerance: 0.0
  max_capability_drop: 0.02
  require:
    - {metric: over_refusal_rate, direction: decrease, min_delta: 0.05}
output:
  repo_id: AmareshHebbar/Qwen2.5-1.5B-Narrow
  license: apache-2.0
```

Any prompt source accepts `dataset` (HuggingFace), `file` (`.txt`, `.jsonl`, `.json`), or an inline `prompts` list, with optional `filter`, `limit`, and `seed`.

Gate options: `harmful_refusal_tolerance`, `max_capability_drop`, `capability_drop_mode` (`absolute` or `relative`), `overrides` (per metric tolerance), and `require` (per metric `increase`, `decrease`, or `not_worse` with `min_delta`).

## Development

```bash
pip install -e ".[dev]"
ruff check src tests
pytest
```

The test suite builds a tiny random Llama with a local tokenizer and runs the full pipeline (narrow sweep, gate, diff, card, patch reload, prune, save and reload) on CPU with no network access.

## Releasing

Releases publish to PyPI through Trusted Publishing from `.github/workflows/publish.yml` (environment `pypi`). Bump `version` in `pyproject.toml`, then:

```bash
git tag v0.1.0
git push origin v0.1.0
```

The workflow runs the tests, checks that the tag matches the package version, builds, and publishes.

## Roadmap

Phase 0 (this release): package, recipes, eval spine, gate, diff, cards, publishing, `narrow`, `prune`, `merge`.

Next: cross architecture `narrow` baselines, prune plus distillation healing, depth upscaling, reasoning conversion, speculative draft models, long context extension, Indic tokenizer extension, multimodal adapters, sparse autoencoders, and a leaderboard Space reading every run's results.

## Scope

The `narrow` method is released for interpretability and evaluation research on over refusal. The release gate blocks any artifact whose refusal on harmful prompts falls below the base model, for every method.

## References

Arditi et al., 2024. Refusal in Language Models Is Mediated by a Single Direction.
Gromov et al., 2024. The Unreasonable Ineffectiveness of the Deeper Layers.
Men et al., 2024. ShortGPT: Layers in Large Language Models are More Redundant Than You Expect.

## License

Apache 2.0
