Metadata-Version: 2.4
Name: targgen
Version: 0.0.1
Summary: 
Author: Trail of Bits
Author-email: Trail of Bits <michele.armillotta@trailofbits.com>
License-Expression: Apache-2.0
License-File: LICENSE
Classifier: Programming Language :: Python :: 3
Requires-Dist: fuzzprep==0.1.0
Requires-Dist: pyyaml>=6
Requires-Python: >=3.13
Project-URL: Homepage, https://pypi.org/project/targgen/
Project-URL: Documentation, https://github.com/trailofbits/targgen/tree/main/docs
Project-URL: Issues, https://github.com/trailofbits/targgen/issues
Project-URL: Source, https://github.com/trailofbits/targgen
Description-Content-Type: text/markdown

# Targgen

**An agentic libFuzzer harnessing tool.**

Targgen pulls in coding agents with different harnessing strategies to generate
fuzzing harnesses for C/C++ libraries.

**Docs:** [strategies](docs/strategies.md) · [coordinator loop](docs/pipeline.md) ·
[config-fuzzing](docs/config-fuzzing.md) · [honesty gates](docs/honesty.md) ·
[CLI & arena spec](docs/cli-reference.md) · [architecture](docs/architecture.md) ·
[security & trust boundary](SECURITY.md)

---

## What a run does

1. **Bootstrap.** Get the source, build the library with coverage and ASan instrumentation,
   and fetch seed corpora. For an OSS-Fuzz project it also fetches the committed harnesses and the
   Fuzz-Introspector coverage snapshot.
2. **Pull in a strategy.** A *strategy* answers "what is worth harnessing?" — for example the
   functions production fuzzing never covers, or the functions that changed most recently. An agent
   reads the strategy's candidate pool and names the targets.
3. **Generate.** One agent per target writes a harness and ensures it reaches the desired
   functionality.
4. **Validity Checks.** A deterministic check replays fixed inputs against the harness. If a fixed
   input reproduces most of the harness's coverage, the harness is not effectively using
   fuzzer-generated data and is subsequently rejected.
5. **Fuzz.** Admitted harnesses go into a shared pool of fuzz slots. A harness whose branches
   another harness already covers is pruned.
6. **Stop and report.** The run ends on the first budget dimension that trips (time or token usage).

---

## Harnessing strategies

Strategies are the core driver behind how harnesses are generated.
List a set of them in the arena's `strategies:`; the coordinator runs them
concurrently and pulls each in again as coverage plateaus.

| Strategy | Selects from | What it harnesses | Needs |
|---|---|---|---|
| [`uncovered`](docs/strategies.md#uncovered) | functions | Uncovered, high-yield **buried functions** from the Fuzz-Introspector snapshot. | An OSS-Fuzz project. On a raw repo the coordinator substitutes `uncovered-from-run`, the same selector over this run's own zero-coverage functions. |
| [`high-value`](docs/strategies.md#high-value) | functions | High-value entry points reasoned **straight from the source**. | — |
| [`git-churn`](docs/strategies.md#git-churn) | functions | Functions with the largest **recent, recency-weighted churn** in Git history. | Readable Git history. |
| [`api-sequencing`](docs/strategies.md#api-sequencing) | sequences | A fixed **producer→consumer call sequence** around one state-carrying handle (the fuzzer supplies data, not order). | The CodeQL CLI. |
| [`harness-variant`](docs/strategies.md#harness-variant) | harnesses | **Config variants** of committed + kept harnesses — vary one axis to unlock new coverage. Always runs in the loop, listed or not. | — |
| [`config`](docs/config-fuzzing.md) | functions | Identify impactful build configurations for the target. | A `config:` block. |
| [`user`](docs/strategies.md#user) | functions | Exactly the functions **you name** in `functions:`. | — |

**Config** is a strategy aligned around pulling in an agent to identify and build non-traditional configurations for the target. `config` works in tandem with other harnessing strategies
to actually produce the fuzzing harness after the build. See [config-fuzzing](docs/config-fuzzing.md).

---

## Requirements

Install these before the first run.

- **Python ≥ 3.13** and [`uv`](https://docs.astral.sh/uv/).
- **LLVM toolchain** on `PATH`: `clang`, `llvm-cov`, `llvm-profdata`.
- The **`claude` CLI**, logged in, on `PATH`. Every agent stage is a `claude` invocation, so run
  `claude` once by hand first and complete the login.
- **CodeQL** — only for the `api-sequencing` strategy. Targgen finds it through `$TARGGEN_CODEQL`,
  `PATH`, or `./codeql/codeql`. Without it, a run that lists `api-sequencing` stops at the analyze
  step.

---

## Disclaimers

**Isolation — run Targgen in a container or a disposable VM.** It compiles and executes
LLM-written C on the host, and its agents hold `Bash`. There is no sandbox. See
[SECURITY.md](SECURITY.md) for the trust boundary, what Targgen writes outside its workspace,
and what leaves the machine.

**Cost — you are responsible for your own API usage.** Every stage of a run invokes the `claude`
CLI, and Targgen applies no spending limit unless you ask for one. Cap a run with
`budget.token_ceiling_usd`, keep `budget.wall_clock_seconds` short on a first run, and monitor
your own usage.

---

## Quickstart

### 1. Set up a source checkout (optional)

```sh
git clone https://github.com/trailofbits/targgen.git
cd targgen && uv sync
```

### 2. Write an `arena.yaml`

Three things are required: a **source**, a **strategy set**, and a **wall-clock budget**.

```yaml
oss_fuzz_project: cjson                  # source: an OSS-Fuzz project name...
strategies: [high-value, harness-variant]  # ...what to harness
budget: { wall_clock_seconds: 3600 }     # ...and when to stop
```

That file is a complete campaign. Everything else has a default. The
[next section](#writing-an-arenayaml) covers the other keys you may want to add.

### 3. Run it

If you cloned the repository, run:

```sh
uv run targgen arena -a arena.yaml --out ./run-bundle
```

To run the published package without cloning, use `uvx` from the directory containing
`arena.yaml`:

```sh
uvx targgen arena -a arena.yaml --out ./run-bundle
```

### 4. Read the results

`--out` gets `report.json`, `report.md`, and a portable `bundle/`. The report holds per-function
coverage, the run's duration and cost, and — for an OSS-Fuzz project — the coverage delta against the committed harnesses.

---

## Writing an `arena.yaml`

### Required keys

| Key | Value | Rule |
|---|---|---|
| `oss_fuzz_project:` **or** `repo_url:` | string | Set at least one. `oss_fuzz_project` gives you the OSS-Fuzz integration, the FI snapshot, and a coverage baseline to compare against. `repo_url` clones any repo directly. Set both to build a raw repo with the OSS-Fuzz integration as a build aid. |
| `strategies:` | list of strings | A non-empty list from the [table above](#harnessing-strategies). An unknown name is an error that lists the valid ones. |
| `budget:` | mapping, with `wall_clock_seconds` | Discovery stops after its rounds, but the fuzz slots keep taking bonus time on kept harnesses, so only a clock ends a run. |

### The keys you are most likely to add

The **Value** column is what the YAML value must be. Targgen rejects the wrong shape at load time —
`strategies: uncovered` fails with "must be a list of strings, got a string" rather than starting a
run that harnesses nothing.

| Key | Value | Default |
|---|---|---|---|
| `model` | string | `sonnet` |
| `effort` | `low`\|`medium`\|`high`\|`xhigh`\|`max` | CLI default |
| `run` | string | `arena` |
| `budget.token_ceiling_usd` | number | none — the run has no spending ceiling until you set one |
| `budget.follow_up_pull_ins` | integer | `6` |
| `concurrency` | mapping of 3 integers | `{max_strategies: 1, max_gen_agents: 2, max_fuzz_slots: 5}` |
| `target_count` | integer | `5` |
| `config` | integer **or** mapping | off |
| `functions` | list of strings | `[]` |
| `judge` | boolean | `false` |
| `configure_args` | list of strings | `[]` |

Every key, its value type, and the sub-keys of `budget`, `concurrency`, and `config` are in the
[arena spec](docs/cli-reference.md#the-arenayaml-spec).

---

## Where the output goes

Two places:

- **`--out DIR`** — the report and the portable bundle. This is what you read and share.
- **The workspace** — everything else: the source checkout, the instrumented build, agent logs,
  corpora, and coordinator state. It defaults to `~/.targgen` and follows `$TARGGEN_WORKSPACE` or
  `-w DIR`. One project's artifacts live under `projects/<project>/`, and each run forks its own
  outputs under `runs/<run>/`, so two runs never clobber each other.

---

## Resuming

If a run is interrupted, pick it up from the state on disk:

```sh
uv run targgen arena -a arena.yaml --resume             # reload state + continue the campaign
uv run targgen arena -a arena.yaml --resume-fuzz-only   # reload + only fuzz existing harnesses
```

Both re-fuzz every non-discarded harness from its saved corpus; `--resume` also keeps pulling in new
strategy work. See [resuming](docs/pipeline.md#resuming).

## Bootstrap primitives (debugging)

The steps the coordinator runs internally are also callable standalone:

```sh
uv run targgen build     -a arena.yaml   # acquire source + build the instrumented library
uv run targgen analyze   -a arena.yaml   # run the strategies' selection-data providers
uv run targgen bootstrap -a arena.yaml   # build + config discovery + analyze
uv run targgen seeds     -a arena.yaml   # collect the seed corpus on its own
```
---

## Contributing

Issues and pull requests are welcome.

```sh
make dev     # sync dev deps + install the prek hooks
make lint    # ruff format --check, ruff check, ty check
make test    # unit tests under test/
make format  # ruff format + ruff check --fix
```

Work on a branch and keep a PR to one logical change. Three checks run on every PR: `make lint`,
`make test`, and `make integration`. That last one crosses the subprocess boundary the unit suite
stops at — it compiles and fuzzes a tiny C library with the real toolchain and installs a built
wheel — so it needs `clang`, `llvm-cov`, and `llvm-profdata` to run locally.

Two things worth knowing before you start:

- **Tuning agent behaviour is usually a prompt edit, not a code change.** The skills under
  `src/targgen/skills/` are meant to be read and edited.
- **Adding a strategy or a data provider** has a checklist:
  [extending Targgen](docs/architecture.md#extending-targgen). Internals:
  [architecture](docs/architecture.md#internals).

Please report security issues through [SECURITY.md](SECURITY.md) rather than a public issue.

## License

Targgen is licensed under [Apache-2.0](LICENSE).
