Metadata-Version: 2.4
Name: llmtrafficlens
Version: 0.11.2
Summary: Profile LLM gateway logs for prefix-cache reuse potential, and export anonymized replayable traces
Project-URL: Homepage, https://github.com/GMISWE/llmtrafficlens
Author: Jason Zhu
License-Expression: Apache-2.0
License-File: LICENSE
Keywords: benchmark,kv-cache,llm,mooncake,prefix-cache,trace,traffic-analysis,workload-characterization
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: System :: Benchmark
Requires-Python: >=3.10
Description-Content-Type: text/markdown

# llmtrafficlens

[![PyPI](https://img.shields.io/pypi/v/llmtrafficlens.svg)](https://pypi.org/project/llmtrafficlens/)
[![Python](https://img.shields.io/pypi/pyversions/llmtrafficlens.svg)](https://pypi.org/project/llmtrafficlens/)
[![License](https://img.shields.io/pypi/l/llmtrafficlens.svg)](LICENSE)

Analyze an LLM gateway's request log for prefix-cache reuse, then export
the log as an anonymized trace that standard benchmark tools can replay.

Reproduce this in a minute on the public Mooncake conversation trace
(`curl -LO https://raw.githubusercontent.com/kvcache-ai/Mooncake/main/FAST25-release/traces/conversation_trace.jsonl`):

```console
$ llmtrafficlens profile conversation_trace.jsonl --format mooncake -o report
12,031 requests · unknown
input p50 6,909 tok · output p50 350 tok · stream 0% · 3.4 req/s

hit rate by cache size  (steady state: last 50% of requests)
      16K tokens    3.6%
      64K tokens    4.5%
     256K tokens    4.5%
       1M tokens    5.6%
       4M tokens   18.8%
      16M tokens   35.6%
       unbounded   39.8%  <- ceiling

session identifiers  0.0% of requests  (affinity routing cannot reach this reuse)
top prefix by volume 512 tok x 12,031 reqs = 5.3% of reusable volume
top prefix by count  512 tok x 12,031 reqs = 5.3% of reusable volume

wrote report.json, report.html
```

`report.html` carries the full leaderboards and distributions;
`report.json` is the same for scripts. That 39.8% ceiling is what the
official `kvcache-simulator` reports for this trace too.

Two commands:

- **`profile`** — how much a prefix cache could save on this traffic, and
  how much cache memory that takes.
- **`export`** — the same log as a Mooncake- or bailian-format trace, with
  all text removed.

## Background

An engine reads the whole prompt before generating anything, and that
work is pure recomputation whenever two requests begin with the same text.
A **prefix cache** keeps the intermediate state around so the next request
starting the same way skips it. Three things decide whether that pays off,
and they are what this tool reports:

- **The traffic.** No shared leading text, nothing to reuse. A property of
  your workload, not your setup.
- **Cache memory.** KV state is bulky, so a cache holds a bounded number
  of tokens and evicts the rest — with diminishing returns you can measure
  instead of guess.
- **Routing.** With several workers, a request only hits if it reaches the
  one holding its prefix. That is what prefix-aware routing does; session
  affinity is the cheaper alternative, and only works when requests carry
  a session identifier.

Reuse is counted in fixed-size **blocks** (16 tokens) because engines
cache at block granularity, not per character.

## Install

```bash
pip install llmtrafficlens
```

No runtime dependencies. Python ≥ 3.10.

## profile

```bash
llmtrafficlens profile gateway.csv -o report
```

**Input.** A CSV with a `request_json` column holding the OpenAI-format
request body; the repository has a synthetic one to try at
`examples/sample_gateway.csv`. These columns are used when present: `response_json`,
`model_name`, `status_code`, and a timestamp column (`timestamp`, `ts`,
`created_at`, ...; epoch or ISO-8601, or set `--ts-column`). The model
name selects which chat-template layout to assume; see
[What counts as the prompt?](#what-counts-as-the-prompt). Mooncake and
qwen-bailian traces are accepted as input too
(`--format mooncake|bailian`).

**Output.** A summary on stdout (shown above), plus `report.html` and
`report.json` containing:

- hit rate at each cache size, ending with the unbounded case — the
  ceiling, and split by what could supply each hit: text shared between
  unrelated conversations, or an earlier turn of the same one. Measured
  over the last half of the log, the first half filling the cache;
  `--warmup 0` measures the whole log from cold instead;
- top prefixes by reuse count, and separately by reusable token volume;
- share of requests carrying a session identifier;
- input/output token distributions, streaming ratio, model mix, QPS.

| option | effect |
|---|---|
| `--warmup F` | leading fraction of requests that fills the cache uncounted (default 0.5) |
| `--model-family F` | whose chat template decides tool placement (default: per request, from the model name) |
| `--tools-position P` | place tool definitions explicitly, overriding the family |
| `--block-size N` | tokens per hash block (default 16) |
| `--ts-column C` | which CSV column holds timestamps |
| `--kvcache-sim ID` | also run kvcache-simulator, see below |

**How to read it.**

| reading | what it decides |
|---|---|
| **ceiling** | whether to bother at all. It is the share of input tokens servable from cache given unlimited memory and perfect routing, so at a few percent no engineering makes this worthwhile |
| **capacity curve** | how much memory the ceiling costs. Where it flattens, buying more stops helping |
| **session-id share** | whether the cheap option suffices. Most requests carrying one means pinning sessions to workers captures the reuse; none carrying one (common for API traffic) leaves prefix matching as the only way |
| **the two leaderboards** | which prefixes must stay resident. The token-volume ranking is the one that determines savings: a short prefix reused constantly contributes almost nothing, a long one reused a few dozen times can be most of it |

### Handing off to kvcache-simulator

Our curve is one LRU replay per capacity. kvcache.ai's simulator does more
on the same input — FIFO, LRU and Belady-optimal side by side, per-model
byte accounting, a C++ core — so `profile` can run it on the same trace
and fold the result in:

```bash
pip install kvcache-simulator
llmtrafficlens profile gateway.csv --kvcache-sim glm-5.2 -o report
```

```console
kvcache-simulator · glm-5.2 · 93 KiB/token
  ceiling 62.1%
        16 GiB  lru 10.4% (1.1x) optimal 31.6% (1.5x)
       128 GiB  lru 44.5% (1.8x) optimal 62.1% (2.6x)
       512 GiB  lru 60.8% (2.6x) optimal 62.1% (2.6x)
```

The gap between the two policies is the part of the shortfall that better
eviction could recover rather than more memory: at 16 GiB above, three
times the hit rate is reachable at the same capacity. The step is skipped
with a note if the simulator is not installed.

## On a production log

The public trace above is chat traffic. A production agent workload looks
different — 2,519 requests through a gateway, tool-calling loops rather
than conversation, each turn resending the whole history:

| cache size | hit rate | across requests | same conversation |
|---|---|---|---|
| 16K tokens | 1.5% | 1.3% | 0.2% |
| 64K | 5.4% | 5.1% | 0.3% |
| 256K | 11.9% | 10.8% | 1.1% |
| 1M | 35.3% | 19.9% | 15.4% |
| 4M | 60.2% | 29.5% | 30.7% |
| unbounded | **62.1%** | 30.9% | 31.2% |

The ceiling splits almost evenly, but the two halves cost very different
amounts of memory. Text shared between unrelated conversations — system
prompts, tool definitions — is touched by hundreds of requests, so LRU
keeps it resident even in a small cache: a third of its potential is
already there at 256K tokens. History resent within one conversation is
touched only by that conversation's later turns, and there is a lot of it
per conversation, so it survives only once the cache can hold many
conversations at once. Below 1M tokens it contributes essentially nothing.

Read that as a sizing rule: a small cache buys you the templates, and the
conversational reuse — half the prize here — needs an order of magnitude
more.

Two findings from the same log worth checking for in your own:

- **No request carried a session identifier.** Half the reuse is
  conversational, so pinning sessions to workers would capture it — but
  nothing in the request says which conversation it belongs to, leaving
  prefix matching as the only way to find out. Having clients send
  `prompt_cache_key` would be cheaper than any routing change.
- **Tool sets changed mid-conversation** in 11% of conversations, one of
  them 191 times. Where a template renders tools near the front, changing
  them breaks the prefix immediately and discards the entire history
  behind it. Keeping the tool array stable for the length of a
  conversation costs nothing and protects the reuse.

## export

```bash
llmtrafficlens export gateway.csv --to mooncake -o trace.jsonl   # or --to bailian
```

A benchmark written by hand, with equally popular prompt groups, reports a
higher hit rate than production reaches, because every group stays warm.
Replaying the real log avoids that, and the export can be shared, because
each request is reduced to one line of structure:

```json
{"timestamp": 1753340000123, "input_length": 1994, "output_length": 117,
 "hash_ids": [4251731047194047120, 8125214104179782736, ...]}
```

`hash_ids` is a salted chained block hash: two requests share their first
N hashes exactly when they share their first N blocks. Replay tools
generate one synthetic block per hash, so the prefix-sharing structure is
preserved while the content is not real.

```bash
# replay at the recorded pace
aiperf profile --custom-dataset-type mooncake_trace --input-file trace.jsonl --fixed-schedule ...
# replay at 2x the pace, for rate sweeps
aiperf profile --custom-dataset-type mooncake_trace --input-file trace.jsonl --synthesis-speedup-ratio 2.0 ...
# or SGLang
python -m sglang.bench_serving --dataset-name mooncake --dataset-path trace.jsonl ...
```

`--to bailian` adds `chat_id`, `parent_chat_id` and `turn` for AIPerf's
`bailian_trace` mode; use it when the log carries session identifiers.
Tell the consumer which block size the export used — AIPerf defaults to
512 for mooncake and 16 for bailian (`--prompt-input-tokens-block-size`).

## Privacy

Processing is local. Reports contain aggregates only. An export contains
lengths, timestamps, anonymized session identifiers and salted block
hashes; prompts cannot be reconstructed from it. The hash key is written
to `.ltl-salt` on first run — keep it unchanged so runs stay comparable,
and treat it as a secret.

## FAQ

### What counts as the prompt?

Everything a chat template renders before generation: the tool
definitions, then each message with its role, name, `tool_call_id`,
content and `tool_calls`. On one production log those extras were 23% of
the request body, so leaving them out inflates the measured reuse
substantially.

Tool definitions go ahead of the conversation, which is where GLM-5.2 and
the DeepSeek family put them. Qwen- and Llama-style templates append them
to the user's system message instead, which lets requests with differing
tool sets keep sharing the system prefix; on the same log that placement
read 6 points higher. We take the conservative one. `reasoning_content` is
excluded, since clients echo it back but engines generally strip it before
prefill.

### How is the hit rate computed?

Requests are replayed in timestamp order against an LRU block cache.
Unbounded capacity gives the "ideal hit rate" of
[KVCache in the Wild (ATC'25)](https://arxiv.org/abs/2506.02634), which is
also the Mooncake simulator's infinite-capacity ceiling; bounded
capacities give that simulator's [capacity
curve](https://kvcache.ai/blog/calculate-kvcache-cache-budge/). LRU is the
default eviction policy in both vLLM (`FreeKVCacheBlockQueue`) and SGLang
(`--radix-eviction-policy lru`), though SGLang evicts radix-tree leaves
rather than flat blocks.

Prefixes are identified by chained per-block hashing (`hash(parent,
block)`, 16-token blocks; 16 is the smallest block vLLM's FlashAttention
backend supports and the granularity it aligns to), following the
[qwen-bailian trace](https://github.com/alibaba-edu/qwen-bailian-usagetraces-anon)
convention.

Three assumptions come with the numbers. The simulation models **one
global cache**, i.e. a single worker or perfect routing, so read the curve
as fleet-level guidance rather than per-worker sizing. An engine's KV pool
also holds in-flight requests, and only the remainder is free for reuse.
And it measures over the **last half of the log**, the first half filling
the cache, so the figure is a steady state rather than one cold replay;
`--warmup 0` gives the cold reading, 37.4% against 39.8% on the Mooncake
trace. The reference simulator makes the same choice, which is why the two
agree.

### How do I convert token counts to gigabytes?

Per-token cost follows the architecture. GLM-5.2 caches two things:

| component | derivation | per token |
|---|---|---|
| MLA latent | `(512 + 64) x 2 bytes x 78 layers` | 87.75 KiB |
| DSA indexer | `128 x 2 bytes x 21 full-indexer layers` | 5.25 KiB |
| | | **93 KiB** |

Only 21 of the 78 layers run a full indexer, the rest reuse the previous
one's selection. Both terms match the `kvcache-simulator` catalog, which
reports 95,232 bytes per token for this model.

Do not scale that figure to another model. Only some architectures carry a
DSA indexer, and the attention term differs as well: Kimi K2.5 is 68.6 KiB
per token, a 32-layer fp16 GQA model 128 KiB. Look yours up with
`kvcache-simulator list-models`.

### Has this been checked against other tools?

Yes, twice, both on the public Mooncake conversation trace with 512-token
blocks, and both pinned by tests in `tests/test_core.py`.

**NVIDIA AIPerf v0.11.0.** `aiperf analyze-trace` reports
`cache_hit_rate: 0.38425808746366197`, which is the unweighted mean of
per-request hit-block ratios. Computing that same definition from our own
block matching yields `0.3842580874636671` — a difference of 5e-15, so the
two implementations agree on every request's prefix match.

**The engine's own rendering.** The strongest check we have, since it
skips our approximations entirely: render each request with the model's
real `chat_template.jinja`, tokenize with its real tokenizer, chunk those
token ids and hash them. On the production log above that pipeline reports
a 61.7% ceiling where ours reports 62.1% — 0.4 points apart, over 51.7M
real tokens versus our 48.8M estimated. Chunking characters rather than
tokens costs almost nothing; getting the tool placement wrong costs 6
points, and dropping tools costs 27.

**kvcache.ai `kvcache-simulator`.** It reports a 39.8% ceiling on this
trace, which our default now reproduces exactly, and an LRU curve of 4.5%
at 30 blocks, 5.5% at 1,937, 17.9% at 7,729 and 34.6% at 30,918 — the same
shape as ours at comparable block counts. Feeding it our own `export`
output works too: it reads the file and lands within 0.1 points of our
figure on the same data.

### What has not been verified?

- **Replay round-trip.** We have not fed an export through AIPerf or
  SGLang and confirmed it replays. The formats match their documented
  schemas, which is not the same thing.
- **Exact tokenization.** Approximate mode hashes character chunks, so
  reported structure is a lower bound on what a real tokenizer would match:
  formatting differences that tokenize identically read as distinct
  prefixes here.
- **Landscape claims.** The statement below that no tool goes from a raw
  gateway log to a hashed trace comes from a survey of vendor
  documentation in July 2026. That is an absence of evidence, not a vendor
  denial.

### Why not just use AIPerf or the Mooncake simulator?

| tool | takes | gives |
|---|---|---|
| **llmtrafficlens** | raw gateway log | hashed trace, reuse structure, session-id presence |
| kvcache-simulator | hashed trace | FIFO/LRU/Optimal, per-model bytes, speedup |
| AIPerf `analyze-trace` | hashed trace | length distributions, cache analysis |
| AIPerf `profile` | hashed trace, SageMaker capture | replay, rate control, latency |
| vLLM / SGLang benchmarks | synthetic or ShareGPT | engine-side prefix-cache performance |

If you already have a hashed trace, use them. `kvcache-simulator` does
more than we do on that input: per-model byte accounting across dozens of
models, three eviction policies side by side, a C++ replay core. AIPerf is
a full benchmarking harness that actually drives an endpoint.

They start from a line like
`{"timestamp": 0, "input_length": 6909, "hash_ids": [1234, 5678, ...]}`.
A gateway log is not that — it is prompts. Turning one into the other
means extracting the prefill text, chunking it, salting and chaining the
hashes, and doing it on a machine you control because that step is the one
that touches real prompts. That is the step this tool covers, and its
`export` output is what those tools consume.

Reading the raw log also preserves two things a hashed trace has already
discarded: whether requests carry session identifiers, which decides
whether cheap affinity routing would work at all, and which specific
prefixes carry the reusable volume. Neither analyzer reports those,
because by the time they see the data it is gone.

## License

Apache-2.0
