Metadata-Version: 2.4
Name: barga
Version: 0.1.1
Summary: barga: a system-one decision model for English and Nepali (inference)
Author: Ampixa
License-Expression: Apache-2.0
Project-URL: Homepage, https://ampixa.com/barga/
Project-URL: Model, https://huggingface.co/ampixa/barga
Project-URL: Source, https://github.com/Ampixa/barga-py
Project-URL: Issues, https://github.com/Ampixa/barga-py/issues
Keywords: nepali,decision,classification,multiple-choice,nlp,modernbert
Classifier: Programming Language :: Python :: 3
Classifier: Operating System :: OS Independent
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: Natural Language :: English
Classifier: Natural Language :: Nepali
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
License-File: NOTICE
Requires-Dist: torch>=2.6
Requires-Dist: transformers<5.19,>=5.17
Requires-Dist: huggingface_hub>=0.30
Requires-Dist: numpy>=1.26
Provides-Extra: onnx
Requires-Dist: onnxruntime>=1.20; extra == "onnx"
Provides-Extra: test
Requires-Dist: pytest>=8; extra == "test"
Dynamic: license-file

# barga

**barga** (Nepali बर्ग, "category") is a small *system-one* decision model for English and Nepali. You give it a
**state** (a conversation, a document, a game position, a record) and one or more **questions**, each with its own
candidate **options**. It returns a probability for every option, in one fast forward pass. It does not generate
text.

This package is the inference code. The weights are on the Hugging Face Hub at
[`ampixa/barga`](https://huggingface.co/ampixa/barga), and the model card has the evaluation and limitations.

## Install

```bash
pip install barga               # PyTorch backend
pip install "barga[onnx]"       # optional onnxruntime backend for exported ONNX bundles
```

Requires Python ≥ 3.10, PyTorch ≥ 2.6 and transformers 5.17 or 5.18. Those transformers versions are the ones
verified to reproduce the reference predictions exactly; see "Verifying a bundle" below.

## Quickstart

```python
import barga

model = barga.load("ampixa/barga")          # Hub repo id, or a local bundle directory; device="cuda" for a GPU

state = ("Caller: Hello, I want to book a check-up for my dog tomorrow morning.\n"
         "Agent: Sure. We have 9:30 or 11:00 tomorrow.\n"
         "Caller: 9:30 is fine.")

decisions = model.decide(state, [
    {"id": "intent", "text": "What does the caller want?",
     "options": [{"id": "book", "description": "Book an appointment"},
                 {"id": "cancel", "description": "Cancel an appointment"},
                 {"id": "info", "description": "Only ask for information"}]},
    {"id": "slot", "text": "Which time did the caller choose?",
     "options": [{"id": "t0930", "description": "9:30 tomorrow"},
                 {"id": "t1100", "description": "11:00 tomorrow"},
                 {"id": "none", "description": "No time chosen yet"}]},
    {"id": "intent_ne", "text": "कल गर्नेले के चाहनुहुन्छ?",
     "options": [{"id": "book", "description": "अपोइन्टमेन्ट बुक गर्न"},
                 {"id": "cancel", "description": "अपोइन्टमेन्ट रद्द गर्न"},
                 {"id": "info", "description": "जानकारी मात्र सोध्न"}]},
])
for d in decisions:
    print(d.question_id, d.choice, d.probs)
```

Output from the released weights (probabilities rounded here):

```
intent    book   {'book': 0.989, 'cancel': 0.009, 'info': 0.003}
slot      t0930  {'t0930': 0.989, 't1100': 0.010, 'none': 0.001}
intent_ne book   {'book': 0.983, 'cancel': 0.002, 'info': 0.015}
```

The state is encoded once and shared by all of its questions, so asking several questions about the same state
costs much less than asking them one at a time.

## Questions and options

Each question is a dict:

| field | required | meaning |
|---|---|---|
| `id` | yes | your id for the question; returned as `question_id` |
| `text` | yes | the question, in English or Nepali (Devanagari or romanized) |
| `options` | yes | 2–20 options, each `{"id", "description"}`; `description` is what the model reads |
| `type` | no | `"choice"` (default), `"boolean"` (exactly two options with `value` `True`/`False`) or `"ordinal"` (numeric `value`s) |
| `rubric` | no | extra instructions read with the question, e.g. what each level of an ordinal scale means |

Option ids are never shown to the model. Only the question text, the rubric and each option's description are.
Ordinal questions also return `expected_value` (the probability-weighted option value).

The model reads at most 1,024 state tokens (when a state is longer, barga keeps its beginning and its most recent
end) and 768 tokens for a question with all of its options. A question that does not fit is rejected with
`barga.SerializationError`; options are never dropped or cut.

See [docs/RECORDS.md](docs/RECORDS.md) for the full input format and how to score canonical barga records with
`decide_record`.

## Reading the output

`decide()` returns one `barga.Decision` per question, in order:

- `choice`: the most probable option id
- `probs`: option id → probability, in your option order; sums to 1
- `expected_value`: for ordinal questions

Probabilities compare options *within* a question. A low top probability (say below 0.6) means the model is unsure,
which is often right for questions the state does not answer. Calibrate a threshold on your own data before acting
on it automatically.

## Verifying a bundle

Every load checks the sha256 of every file against the bundle's `bundle.json` and refuses a mismatch. The package's
parity test reproduces barga's reference evaluation bit for bit on real records:

```bash
pip install "barga[test]"
BARGA_TEST_BUNDLE=/path/to/bundle BARGA_TEST_PARITY=/path/to/parity-dir pytest tests/test_bundle_parity.py
```

## Licence

Apache-2.0 (see LICENSE and NOTICE). The weights are fine-tuned from Julia-1 (Supersonic Labs, Apache-2.0), which is
built on mmBERT-small (JHU CLSP, MIT).
