Metadata-Version: 2.4
Name: thx01
Version: 1.0.0
Summary: THX-01: calibrated, non-autoregressive typed decisions in one forward pass (100+ languages, Azerbaijani-first)
Author: Elturan Ahmadbayli
Author-email: Farid Aghayev <farid.a@hal-x.ai>
License: Apache-2.0
Project-URL: Homepage, https://huggingface.co/doofz/THX-01
Project-URL: Documentation, https://huggingface.co/doofz/THX-01#quickstart
Project-URL: API docs, https://api.hal-x.ai/docs/thx-01/
Keywords: decision model,classification,routing,calibration,multilingual,azerbaijani,extraction,llm routing,typesafe
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Programming Language :: Python :: 3
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Text Processing :: Linguistic
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
License-File: NOTICE
Requires-Dist: torch>=2.0.0
Requires-Dist: transformers>=4.48.0
Requires-Dist: safetensors>=0.4.0
Requires-Dist: huggingface_hub>=0.20.0
Requires-Dist: numpy>=1.20.0
Provides-Extra: server
Requires-Dist: fastapi>=0.110; extra == "server"
Requires-Dist: uvicorn>=0.29; extra == "server"
Provides-Extra: doom
Requires-Dist: vizdoom>=1.2; extra == "doom"
Requires-Dist: pillow; extra == "doom"
Requires-Dist: fastapi>=0.110; extra == "doom"
Requires-Dist: uvicorn>=0.29; extra == "doom"
Dynamic: license-file

# THX-01

THX-01 is a non-autoregressive, multilingual decision model developed by HAL-X AI. Given a state (a message, ticket, e-mail,
document, JSON record or agent trace) and one or more typed questions written in natural language, it returns a calibrated
answer to every question in a single forward pass of about 10 ms on one GPU.

THX-01 is trained with large-scale Reinforcement Learning for Calibrated Decisions (RLCD): the reward is a strictly proper
scoring rule, so reporting honest probabilities is the only way to maximise it. Beyond choosing among options, THX-01 can
return numbers stated in a document, verbatim excerpts and supporting citations, through the same interface.

| | |
|---|---|
| Parameters | 322M (307M encoder, 15M decision head) |
| Encoder | mmBERT-base, 22 layers, 256k vocabulary |
| Input | up to 1,024 tokens per question (question, options and state); longer states are truncated |
| Post-training | eight large-scale RLCD stages, about 2.1 million training decisions |
| Languages | post-trained in 18 languages with emphasis on Azerbaijani; 100+ supported |
| Latency | about 10 ms per request, about 1 ms per decision when batched |
| License | Apache 2.0 |

## Contents

1. [Question types](https://huggingface.co/doofz/THX-01#question-types)
2. [Quickstart](https://huggingface.co/doofz/THX-01#quickstart)
3. [Benchmarks](https://huggingface.co/doofz/THX-01#benchmarks)
4. [Method](https://huggingface.co/doofz/THX-01#method)
5. [Training](https://huggingface.co/doofz/THX-01#training)
6. [Languages](https://huggingface.co/doofz/THX-01#languages)
7. [REST API](https://huggingface.co/doofz/THX-01#rest-api)
8. [Limitations](https://huggingface.co/doofz/THX-01#limitations)
9. [Citation](https://huggingface.co/doofz/THX-01#citation)

## Question types

| type | returns | criteria |
|---|---|---|
| `choice` | one of N options with a probability for each | `{"key": "description", ...}` or a list |
| `noul` | P(yes) | optional |
| `score` | an ordinal level and its distribution | ordered list of levels |
| `number` | a numeric value stated in the document, or `null` if it is not stated | optional `min`, `max`, `precision` |
| `excerpt` | a verbatim span of the document with its character offsets | none |
| `"cite": true` (any question) | the parts of the document that support the answer, with probabilities | none |

`number`, `excerpt` and citations are native THX-01 capabilities. They are not part of the TypeSafe question types
(`choice`, `noul`, `score`), where they have to be emulated in the service layer with repeated `choice` calls.
THX-01 also answers such emulation calls (value-range buckets and document chunks) through its native lookup, so existing
service layers work unchanged.

`number` returns values exactly as written in the document, normalised (`1,2 mln` becomes `1200000`). It performs no unit
or currency conversion: a question about kilometres when the document states miles, or about euros when it states US
dollars, is answered with `null`. `excerpt` cuts its answer out of the document, so it cannot contain invented text.

## Quickstart

```python
import thx01

agent = thx01.load("doofz/THX-01")              # GPU if available

doc = ("Northwind Corp. reported Q3 2026 results. Revenue for the quarter was $48.3 million, up 14%. "
       "Net income was $6.2 million, compared with $4.1 million a year ago. "
       '"We will open two offices in Baku," said CEO Laura Chen.')

result = agent.decide(doc, {
    "kind":    {"type": "choice", "question": "What is this document?",
                "criteria": {"earnings": "earnings report", "complaint": "customer complaint", "invoice": "invoice"}},
    "revenue": {"type": "number", "question": "What was the revenue, in US dollars?", "cite": True},
    "quote":   {"type": "excerpt", "question": "What did the CEO say?"},
})
# kind -> earnings, revenue -> 48300000 (+ the supporting sentence),
# quote -> verbatim span starting "We will open two offices in Baku"
```

Install from PyPI (the weights download from this repository on first use):

```bash
pip install thx01                 # library
pip install "thx01[server]"       # plus the REST server
```

## Benchmarks

### Support-ticket classification

Fifteen categories (billing, refund, login, technical bug, outage, delivery, order change, product question, account change,
security and fraud, data privacy, complaint, feature request, integration and API, contract and legal), four test sets with
2,843 tickets in Azerbaijani, Russian, English and Turkish: Clean, Corrupted (transliteration, removed diacritics, typos,
noise), Messy (written messy on purpose) and Independent (written by GPT-6-Luna, never used for training).

<img src="https://huggingface.co/doofz/THX-01/resolve/main/assets/thx01_ticket_bars.png" width="100%" />

| model | Clean | Corrupted | Messy | Independent | Avg | ECE | latency |
|---|---|---|---|---|---|---|---|
| **THX-01** | 99.2 | 97.7 | 98.8 | 97.9 | 98.4 | 0.003 | ~10 ms |
| Claude Sonnet 5.5 † | 98.5 | 98.0 | 98.5 | 99.0 | 98.5 | – | 1.5 s |
| Wahoo 1.5 | 99.8 | 96.3 | 98.0 | 97.2 | 97.8 | – | 145 ms |
| GPT-6-Luna | 98.6 | 97.4 | 96.7 | 98.2 | 97.7 | – | 1.9 s |
| TypeSafe Jev 1.13 | 99.2 | 95.0 | 96.2 | 99.0 | 97.4 | 0.007 | 331 ms |
| Kev-4B | 96.1 | 88.4 | 90.8 | 95.8 | 92.8 | 0.202 | 830 ms |

† Evaluated on a stratified subset of 200 tickets per set. ECE is measured on the Independent set; LLMs return no
probabilities. THX-01 latency on one GPU; other models through their APIs.

<img src="https://huggingface.co/doofz/THX-01/resolve/main/assets/robustness.png" width="100%" />

### Extraction and citation

Held-out multilingual documents (earnings reports, invoices, customs declarations, contracts, news, e-mails): 1,824 number
questions and 1,357 excerpt questions. Number tasks report exact-value accuracy, excerpt tasks token F1, citation the
accuracy of the top supporting part.

<img src="https://huggingface.co/doofz/THX-01/resolve/main/assets/thx01_extraction.png" width="100%" />

| task | previous checkpoint | THX-01 |
|---|---|---|
| Number lookup | 69.6 | **93.4** |
| Number via range buckets (service-layer emulation) | 0.0 | **95.1** |
| Excerpt extraction (F1) | 20.1 | **84.1** |
| Excerpt via document chunks (F1) | 8.9 | **83.8** |
| Reference citation | 12.7 | **94.0** |

On number lookup over the same documents, TypeSafe Jev 1.13 reaches 96.6%.

### Calibration and selective automation

<img src="https://huggingface.co/doofz/THX-01/resolve/main/assets/reliability.png" width="100%" />

Because confidence is calibrated, a threshold turns it into an operating policy: THX-01 handles about three quarters of all
tickets automatically without a single error on any of the four sets, and 90% of tickets at an accuracy of at least 99.5%.

<p float="left">
  <img src="https://huggingface.co/doofz/THX-01/resolve/main/assets/coverage.png" width="49%" />
  <img src="https://huggingface.co/doofz/THX-01/resolve/main/assets/latency.png" width="49%" />
</p>

### Multilingual decision suite

Thirty-nine held-out tasks: LLM routing, held-out synthetic schemas in 16 languages, MASSIVE scenarios and intents,
SIB-200 topic classification in ten languages, AG News, Banking77, SMS spam, DAIR Emotion and Azerbaijani app reviews.
Mean accuracy rises from 58.3% at initialisation to **84.2%**, and mean ECE falls from
0.204 to **0.066**. Hand-written LLM routing reaches 100%.

<img src="https://huggingface.co/doofz/THX-01/resolve/main/assets/suite.png" width="100%" />

<details>
<summary>All 39 tasks</summary>

| task | initialisation | THX-01 | ECE |
|---|---|---|---|
| LLM routing (hand-written) | 30.9 | 100.0 | 0.060 |
| Synthetic routing (held-out schemas) | 33.2 | 95.6 | 0.020 |
| Synthetic decisions (held-out schemas) | 53.4 | 87.8 | 0.025 |
| Synthetic guard (held-out schemas) | 66.9 | 96.8 | 0.024 |
| Synthetic taxonomy (held-out schemas) | 36.0 | 90.6 | 0.035 |
| Synthetic emotion (held-out schemas) | 58.6 | 96.4 | 0.051 |
| AG News (4) | 94.1 | 92.3 | 0.018 |
| DAIR Emotion (6) | 43.4 | 44.6 | 0.284 |
| Banking77 (77) | 51.7 | 65.9 | 0.087 |
| SMS spam (yes/no) | 73.5 | 87.3 | 0.048 |
| AZ app-review sentiment | 73.5 | 87.0 | 0.072 |
| MASSIVE-intent@20 (az) | 36.3 | 89.7 | 0.038 |
| Intent yes/no (az) | 65.2 | 95.2 | 0.017 |
| MASSIVE-intent@20 (en) | 61.9 | 93.5 | 0.041 |
| Intent yes/no (en) | 66.7 | 95.3 | 0.017 |
| MASSIVE-scenario (en) | 69.4 | 90.8 | 0.049 |
| MASSIVE-scenario (az) | 41.6 | 88.8 | 0.052 |
| MASSIVE-scenario (ru) | 57.6 | 89.8 | 0.039 |
| MASSIVE-scenario (tr) | 50.2 | 87.4 | 0.055 |
| MASSIVE-scenario (de) | 56.4 | 88.2 | 0.050 |
| MASSIVE-scenario (fr) | 59.8 | 90.4 | 0.044 |
| MASSIVE-scenario (es) | 55.0 | 87.8 | 0.040 |
| MASSIVE-scenario (ar) | 43.8 | 82.2 | 0.050 |
| MASSIVE-scenario (hi) | 46.2 | 85.4 | 0.037 |
| MASSIVE-scenario (zh-CN) | 61.0 | 87.2 | 0.060 |
| MASSIVE-scenario (fa) | 46.6 | 89.0 | 0.052 |
| MASSIVE-scenario (ka) | 16.8 | 78.0 | 0.060 |
| MASSIVE-scenario (ja) | 59.2 | 91.8 | 0.043 |
| MASSIVE-scenario (ko) | 48.8 | 86.4 | 0.039 |
| SIB-200 (az) | 67.2 | 73.5 | 0.111 |
| SIB-200 (en) | 78.4 | 78.4 | 0.068 |
| SIB-200 (ru) | 75.5 | 76.0 | 0.082 |
| SIB-200 (tr) | 74.0 | 73.5 | 0.126 |
| SIB-200 (de) | 75.5 | 79.9 | 0.080 |
| SIB-200 (ar) | 73.0 | 76.5 | 0.076 |
| SIB-200 (hi) | 67.2 | 69.6 | 0.132 |
| SIB-200 (zh) | 77.9 | 77.9 | 0.082 |
| SIB-200 (kk) | 68.6 | 68.1 | 0.158 |
| SIB-200 (uz) | 58.3 | 70.1 | 0.145 |

</details>

### Speed

| workload | time on one GPU |
|---|---|
| one question | about 9 ms |
| three questions about the same state | about 10 ms |
| 150 requests batched | about 156 ms |
| peak throughput | about 3,000 decisions per second |

## Method

Every question is serialised as `[CLS] question [MASK] option 1 [MASK] option 2 ... [SEP] state`. The encoder reads the
whole sequence at once; a two-layer decision head scores each option at its own `[MASK]` position, and a softmax with a
temperature fitted per question type and option count gives calibrated probabilities. All questions of a request are
scored in one batched pass. Questions with more than 24 options are decided by a two-round tournament.

### Reinforcement Learning for Calibrated Decisions

Each question is a one-step game: the policy reports a distribution q over the options, the outcome y is revealed, and the
reward is

R(q, t) = sum_k t_k log q_k + 0.5 * (sum_k t_k q_k) / ||q||_2 - 1[ordinal] * RPS(q, t)

a combination of the logarithmic, spherical and ranked-probability scores. Because the reward is strictly proper, its
expectation is maximised only when the reported distribution equals the true one. THX-01 optimises the expected reward with
an exact, zero-variance gradient. Soft targets teach two further behaviours: a uniform target when no option applies (so
the model reports uncertainty instead of a confident wrong answer) and near-miss credit on ordinal scales.

## Training

<img src="https://huggingface.co/doofz/THX-01/resolve/main/assets/training.png" width="100%" />

| stage | focus | decisions per epoch |
|---|---|---|
| S1 | broad multilingual RLCD: synthetic decisions in 16 languages, MASSIVE in 14 languages, Azerbaijani sentiment | 157,912 (x2) |
| S2 | spatial and control decisions | 53,040 |
| S3 | LLM and agent routing | 48,000 |
| S4 | generic skills: bare-key options, identifier lookup, no-good-option honesty | 77,000 |
| S5 | robustness and domain: 600 new business taxonomies, confusable pairs, transliteration and noise, support tickets | 213,961 (x2) |
| S6 | boundary cases between confusable categories, three teacher models | 284,207 |
| S7 | weight interpolation and temperature refit | – |
| S8 | extraction and citation: numbers, excerpts, supporting parts, currency equivalence, new hard decisions, full replay | 493,442 |

Training data combines curated data from the HAL-X data team, public benchmarks' training splits and verified synthetic
data. Synthetic items are generated label-first by teacher models (Wahoo 1.5, GPT-6-Luna, Claude Sonnet 5.5) and kept only
when a blind second pass agrees; test items and their near-duplicates are excluded from training. Optimisation uses 8-bit
AdamW, bf16, length-bucketed batches, random option order and option subsets, and a frozen token-embedding matrix that
preserves the encoder's coverage of more than 1,800 pretraining languages.

<img src="https://huggingface.co/doofz/THX-01/resolve/main/assets/progression.png" width="60%" />

## Languages

Post-training covers Azerbaijani, English, Russian, Turkish, German, French, Spanish, Arabic, Hindi, Chinese, Persian,
Georgian, Japanese, Korean, Ukrainian, Kazakh, Uzbek and Italian, with Azerbaijani the largest language in every synthetic
component. A dedicated curriculum covers the way text is actually typed in the region: Azerbaijani without its letters
(e for the schwa, s for s-cedilla, c for c-cedilla), Russian in Latin transliteration, and code-mixing of Azerbaijani,
Russian and English. Through its encoder THX-01 supports more than 100 languages.

| model | Independent az | ru | en | Corrupted az | ru | en |
|---|---|---|---|---|---|---|
| **THX-01** | 97.1 | 98.3 | 98.3 | 97.9 | 96.6 | 97.9 |
| Claude Sonnet 5.5 † | 98.5 | 100.0 | 98.5 | 97.6 | 100.0 | 97.5 |
| Wahoo 1.5 | 96.2 | 95.8 | 99.6 | 94.5 | 95.0 | 98.9 |
| GPT-6-Luna | 97.9 | 97.5 | 99.2 | 97.2 | 95.0 | 98.6 |
| TypeSafe Jev 1.13 | 98.8 | 98.3 | 100.0 | 92.9 | 93.3 | 98.2 |
| Kev-4B | 94.2 | 95.4 | 97.9 | 83.1 | 89.1 | 94.3 |

## REST API

```bash
pip install "thx01[server]"
THX01_MODEL=doofz/THX-01 THX01_API_KEY=your-key python -m thx01.server --port 8095
```

`POST /v1/decide` (alias `POST /v1/systemone`, TypeSafe-compatible) takes `{"state": ..., "questions": {...}}` and returns
`{"answers": {...}, "latency_ms": ...}`; `POST /v1/decide/batch` (alias `/v1/systemone/batch`) takes `{"items": [...]}`.
The state may be a string or an object with a `document` field; the question text may be given as `instructions` or
`question`. Requests are micro-batched on the GPU.

## Limitations

- Fine-grained intent sets with many near-synonymous labels remain hard: Banking77 (77 intents) reaches 65.9%.
- Emotion with six overlapping classes (DAIR Emotion) reaches 44.6%; the model reports correspondingly low confidence.
- `number` and `excerpt` locate text that is present in the document; they do not perform arithmetic, unit conversion or
  paraphrase.

## Citation

```bibtex
@techreport{thx01_2026,
  title       = {THX-01: Large-Scale Reinforcement Learning for Calibrated Decisions in 100+ Languages},
  author      = {Aghayev, Farid and Ahmadbayli, Elturan},
  institution = {HAL-X AI},
  year        = {2026}
}
```

## License

Apache License 2.0. See [LICENSE](https://huggingface.co/doofz/THX-01/blob/main/LICENSE) and [NOTICE](https://huggingface.co/doofz/THX-01/blob/main/NOTICE).
