Metadata-Version: 2.4
Name: ftgate
Version: 0.1.0
Summary: Regression checks for small fine-tuned models on tool calling: do runtimes send the bytes the model was trained on?
Author: Vitalii Bogachev
License: Apache-2.0
Project-URL: Homepage, https://github.com/talhayme/ftgate
Project-URL: Repository, https://github.com/talhayme/ftgate
Project-URL: Issues, https://github.com/talhayme/ftgate/issues
Keywords: fine-tuning,tool-calling,ollama,llama.cpp,gguf,evaluation,dataset,lint
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Software Development :: Testing
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: transformers>=4.45
Requires-Dist: requests>=2.31
Provides-Extra: dev
Requires-Dist: pytest; extra == "dev"
Dynamic: license-file

# ftgate (working name)

Regression checks for a small model you fine-tuned on tool calling: do the
bytes the runtime sends match the bytes the model was trained on, and does
it still call your tools correctly after export and quantisation?

**Status: `ftgate template`, `ftgate data` and `ftgate tools` work.** The premise was
checked before any tool code was written — the same way
[pr-witness](https://github.com/talhayme/pr-witness) was built.

## What has been measured so far

- **llama-server reproduces the HF training bytes exactly** for a
  tool-calling prompt (Qwen2.5-0.5B-Instruct, 250/250 tokens, text and ids
  identical) — even though the GGUF's embedded template hashes differently
  from HF's. Compare renderings, never hashes. → [`notes/day1-findings.md`](notes/day1-findings.md)
- **Ollama 0.40.2 hands the model a Go struct dump instead of the tool
  schema**: `{get_weather … {object <nil> <nil> [city] {…}}}`. Root cause
  in `template/template.go` (`templateToolFunction` has no `String()`);
  known upstream as ollama/ollama#14601 with open PRs #18318 and #18391.
- **Qwen's official Qwen2.5 GGUFs embed the pre-fix chat template** —
  doubled braces `{{"name": …}}` in the tool-call instruction, fixed on HF
  on 2024-09-19 but never regenerated in the GGUF repos. Ollama's library
  blobs are those files. Ollama never renders it; llama.cpp, vLLM or LM
  Studio serving the file do. **The 0.5B model copies the braces into
  every tool call** (invalid JSON; a strict parser gets 0/35 where the HF
  template gets 22/35); the 1.5B model ignores them. llama-server's parser
  forgives it, vLLM's `json.loads` would not.
  → [`notes/day4-doubled-braces-impact.md`](notes/day4-doubled-braces-impact.md)
  → [`notes/day3-template-command.md`](notes/day3-template-command.md)
- **Ollama prepends the Modelfile SYSTEM** when a request has none — a
  system prompt your training rows may never have had.
- **It costs accuracy.** Same GGUF blob in every arm, temperature 0, 40
  cases: Qwen2.5-0.5B gets exact arguments 39/40 with its own template and
  31/40 with the prompt Ollama builds — every failure an invented
  parameter `amount from_currency`, read off Go's `[amount from_currency
  to_currency]` rendering of `required`. Qwen2.5-1.5B: 40/40 either way on
  this easy set. → [`notes/day2-impact.md`](notes/day2-impact.md)

## Use

```
uv pip install -e .
ftgate template --model Qwen/Qwen2.5-0.5B-Instruct \
                --llama http://localhost:8080 \
                --ollama qwen2.5:0.5b-instruct \
                --dataset train.jsonl --sample 20 --diff
```

`--model` is the HF checkpoint whose tokenizer and chat template the
trainer used; `--dataset` is your own rows (`messages[, tools]` per line).
Each runtime's prompt and token ids are compared with what SFT fed the
model, and every difference is classified: `schema_not_json`,
`injected_system`, `special_prefix`, `structural`, `spacing`, `tokenizer`.
Ollama's column is rendered with Ollama's own `template` package (build
`bin/` once) and cross-checked against the live `prompt_eval_count`.

### `ftgate data` — lint the dataset with the model's own tokenizer and template

```
ftgate data train.jsonl --model Qwen/Qwen2.5-0.5B-Instruct --max-seq-len 2048 \
            --eval eval.jsonl --fail-on error
```

Rows in `messages`, ShareGPT, Alpaca or preference (`chosen`/`rejected`)
shape. Reports what the trainer will silently do to each row: an assistant
turn that renders to zero target tokens, a conversation cut by
`max_seq_len` before the answer (trains nothing) or inside it, a row that
never ends with EOS, a template that raises on the row. Structural checks:
empty or missing assistant target, role order, tool calls against the
declared schema (missing required / unknown arguments / undeclared tool /
`tool_call_id` with no call), `chosen == rejected`, length bias in
preference pairs, exact duplicates, near-duplicates of the eval set.
`--format json` for CI; `--fail-on error|warning` sets the exit code.

### `ftgate tools` — does it still call *your* tools, on the backend it will run on?

```
ftgate tools --tools tools.json --cases cases.jsonl \
             --provider ollama:my-finetune:q4_k_m \
             --provider "openai:m@http://localhost:8080/v1" \
             --min-exact 0.8
ftgate tools --from-dataset train.jsonl --holdout 50 --provider ollama:my-finetune
```

Evaluation runs on [promptfoo](https://www.promptfoo.dev/) (installed
once with `npm install --prefix .pfrunner promptfoo`, or set
`FTGATE_PROMPTFOO`). ftgate writes the config — your tools, your cases
(a file, or rows held out of your own training set), one provider per
arm, temperature 0, one request in flight for local servers — and a
deterministic judge that reports per case: the right tool, every expected
argument present and equal (extra optional arguments allowed and counted),
and two things a generic eval does not separate from "the model was
wrong": **parser lost** — the model emitted a `<tool_call>` the provider
did not surface — and **malformed** — the call's JSON does not parse.
Failure reasons show the produced arguments next to the expected ones.

### In CI, in tests, before commit

```yaml
# GitHub Actions
- uses: talhayme/ftgate@main
  with: { dataset: data/train.jsonl, model: Qwen/Qwen2.5-0.5B-Instruct, max-seq-len: "2048" }
```

```yaml
# .pre-commit-config.yaml
- repo: https://github.com/talhayme/ftgate
  rev: main
  hooks: [{ id: ftgate-data, args: [--fail-on, error] }]
```

```python
# tests/test_data.py
from ftgate.testing import assert_dataset_clean

def test_training_data():
    assert_dataset_clean("data/train.jsonl", model="Qwen/Qwen2.5-0.5B-Instruct", max_seq_len=2048)
```

or, with the plugin that installs alongside the package,
`pytest --ftgate-dataset data/train.jsonl --ftgate-model Qwen/Qwen2.5-0.5B-Instruct`.

## Reproduce the findings

```
uv venv --python 3.12 ~/.venvs/ftgate && uv pip install --python ~/.venvs/ftgate/bin/python torch transformers gguf requests
# llama-server built from llama.cpp with --jinja support; Ollama ≥ 0.40 running on :11434
python scripts/hf_side.py                       # training-side bytes
python scripts/llama_side.py http://localhost:8089
python scripts/ollama_side.py                   # needs bin/ollama-render (see bin/README.md)
python scripts/day2_eval.py qwen2.5:0.5b-instruct   # three-arm accuracy comparison
```

## Plan

`ftgate template` — the three-sided byte comparison as a command.
`ftgate data` — a linter for SFT/DPO/tool-call datasets using the target
model's tokenizer and template. Tool-call evaluation builds on promptfoo's
`tool-call-f1` / `trajectory:tool-args-match` rather than duplicating them.

Apache-2.0.
