Metadata-Version: 2.4
Name: llmtrafficlens
Version: 0.7.0
Summary: Profile LLM gateway logs for prefix-cache reuse potential, and export anonymized replayable traces
Project-URL: Homepage, https://github.com/GMISWE/llmtrafficlens
Author: Jason Zhu
License-Expression: Apache-2.0
License-File: LICENSE
Keywords: benchmark,kv-cache,llm,mooncake,prefix-cache,trace,traffic-analysis,workload-characterization
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: System :: Benchmark
Requires-Python: >=3.10
Description-Content-Type: text/markdown

# llmtrafficlens

[![PyPI](https://img.shields.io/pypi/v/llmtrafficlens.svg)](https://pypi.org/project/llmtrafficlens/)
[![Python](https://img.shields.io/pypi/pyversions/llmtrafficlens.svg)](https://pypi.org/project/llmtrafficlens/)
[![License](https://img.shields.io/pypi/l/llmtrafficlens.svg)](LICENSE)

Analyze an LLM gateway's request log for prefix-cache reuse, then export
the log as an anonymized trace that standard benchmark tools can replay.

On the public Mooncake conversation trace, which the
[Try it](#try-it-on-a-public-trace) section below reproduces verbatim:

```console
$ llmtrafficlens profile conversation_trace.jsonl --format mooncake -o report
12,031 requests · unknown
input p50 6,909 tok · output p50 350 tok · stream 0% · 3.4 req/s

hit rate by cache size  (steady state: last 50% of requests)
      16K tokens    3.6%
      64K tokens    4.5%
     256K tokens    4.5%
       1M tokens    5.6%
       4M tokens   18.8%
      16M tokens   35.6%
       unbounded   39.8%  <- ceiling

session identifiers  0.0% of requests  (affinity routing cannot reach this reuse)
top prefix by volume 512 tok x 12,031 reqs = 5.3% of reusable volume
top prefix by count  512 tok x 12,031 reqs = 5.3% of reusable volume

wrote report.json, report.html
```

`report.html` carries the full leaderboards and distributions;
`report.json` is the same content for scripts.

Two commands:

- **`profile`** — how much a prefix cache could save on this traffic, and
  how much cache memory that takes.
- **`export`** — the same log as a Mooncake- or bailian-format trace, with
  all text removed.

## Background

An engine serves a request in two phases: **prefill** reads the whole
prompt at once, then **decode** emits output tokens one by one. Prefill
cost grows with prompt length, and it is pure recomputation whenever two
requests begin with the same text — the same system prompt, the same
few-shot examples, the same attached document.

A **prefix cache** avoids that: the intermediate state a prompt produces
(its KV cache) stays in memory, and the next request starting with the
same text reuses it instead of recomputing. Whether that pays off depends
on three things, which are exactly what this tool reports:

- **The traffic.** If requests share no leading text, nothing can be
  reused. This is a property of your workload, not of your setup.
- **Cache memory.** KV state is bulky, so a cache holds a bounded number
  of tokens and evicts the rest. More memory, more hits — with
  diminishing returns you can measure instead of guess.
- **Routing.** Across several workers, a request only hits if it reaches
  the worker that already holds its prefix. Sending it there is what
  prefix-aware routing does; the alternative, session affinity, only works
  when requests carry a session identifier.

Reuse is tracked in fixed-size **blocks** (16 tokens by default) because
engines cache at block granularity, not per character.

## Install

```bash
pip install llmtrafficlens
```

No runtime dependencies. Python ≥ 3.10.

## profile

```bash
llmtrafficlens profile gateway.csv -o report
```

**Input.** A CSV with a `request_json` column holding the OpenAI-format
request body. These columns are used when present: `response_json`,
`model_name`, `status_code`, and a timestamp column (`timestamp`, `ts`,
`created_at`, ...; epoch or ISO-8601, or set `--ts-column`). Mooncake and
qwen-bailian traces are accepted as input too
(`--format mooncake|bailian`).

**Output.** A summary on stdout (shown above), plus `report.html` and
`report.json` containing:

- hit rate at each cache size, ending with the unbounded case — the
  ceiling. Measured over the last half of the log, the first half filling
  the cache; `--warmup 0` measures the whole log from cold instead;
- top prefixes by reuse count, and separately by reusable token volume;
- share of requests carrying a session identifier;
- input/output token distributions, streaming ratio, model mix, QPS.

**How to read it.**

*Ceiling* — the share of input tokens that could be served from cache if
memory were unlimited and every request reached the right worker. It
bounds everything else: at a few percent, no amount of engineering makes
prefix caching worthwhile on this traffic.

*Capacity curve* — the hit rate at each cache size, so you can see what
the ceiling costs. Where it flattens is the point past which buying
memory stops helping.

*Session-identifier share* — whether the cheap option is enough. When most
requests carry an identifier, pinning each session to a worker captures
the reuse. When none do (a common case for API traffic), the reuse sits in
prefixes shared between unrelated requests, and only prefix-aware routing
reaches it.

*The two leaderboards* — which prefixes to keep resident. They rank
differently, and the token-volume one is what determines savings: a short
prefix reused very often contributes almost no reusable volume, while a
long prefix reused a few dozen times can account for most of it.

## export

```bash
llmtrafficlens export gateway.csv --to mooncake -o trace.jsonl   # or --to bailian
```

A benchmark written by hand, with equally popular prompt groups, reports a
higher hit rate than production reaches, because every group stays warm.
Replaying the real log avoids that, and the export can be shared, because
each request is reduced to one line of structure:

```json
{"timestamp": 1753340000123, "input_length": 1994, "output_length": 117,
 "hash_ids": [4251731047194047120, 8125214104179782736, ...]}
```

`hash_ids` is a salted chained block hash: two requests share their first
N hashes exactly when they share their first N blocks. Replay tools
generate one synthetic block per hash, so the prefix-sharing structure is
preserved while the content is not real.

```bash
# replay at the recorded pace
aiperf profile --custom-dataset-type mooncake_trace --input-file trace.jsonl --fixed-schedule ...
# replay at 2x the pace, for rate sweeps
aiperf profile --custom-dataset-type mooncake_trace --input-file trace.jsonl --synthesis-speedup-ratio 2.0 ...
# or SGLang
python -m sglang.bench_serving --dataset-name mooncake --dataset-path trace.jsonl ...
```

`--to bailian` adds `chat_id`, `parent_chat_id` and `turn` for AIPerf's
`bailian_trace` mode; use it when the log carries session identifiers.
Tell the consumer which block size the export used — AIPerf defaults to
512 for mooncake and 16 for bailian (`--prompt-input-tokens-block-size`).

## Try it on a public trace

No data of your own needed:

```bash
curl -LO https://raw.githubusercontent.com/kvcache-ai/Mooncake/main/FAST25-release/traces/conversation_trace.jsonl
llmtrafficlens profile conversation_trace.jsonl --format mooncake -o mooncake-report
```

That prints the summary at the top of this README, and reproduces the
39.8% ceiling the official `kvcache-simulator` reports for this trace.
Reading the curve: this workload gains almost nothing below 1M tokens, and
most of its ceiling needs a cache past 16M — which for GLM-5.2 at bf16
(93 KiB/token) is 1.5 TiB, more than device memory alone holds, and the
case tiered KV stores (host memory, SSD) are built for.

Cache size is reported in tokens because bytes are model-specific:
93 KiB/token for GLM-5.2, 68.6 for Kimi K2.5, 128 for a 32-layer fp16 GQA
model. Some architectures also carry a DSA indexer cost on top of
attention. Look yours up with `kvcache-simulator list-models`.

The repository also contains a synthetic sample log (`examples/`, not
shipped in the pip package): 360 requests, three shared system prompts
with skewed popularity, sparse session identifiers, real timestamps. It
reports a 63.0% ceiling, and it shows the two leaderboards disagreeing —
16 tokens × 180 requests is 2.3% of the reusable volume, while 2400
tokens × 30 requests is 55.6% of it.

## Privacy

Processing is local. Reports contain aggregates only. An export contains
lengths, timestamps, anonymized session identifiers and salted block
hashes; prompts cannot be reconstructed from it. The hash key is written
to `.ltl-salt` on first run — keep it unchanged so runs stay comparable,
and treat it as a secret.

## FAQ

### How is the hit rate computed?

Requests are replayed in timestamp order against an LRU block cache.
Unbounded capacity gives the "ideal hit rate" of
[KVCache in the Wild (ATC'25)](https://arxiv.org/abs/2506.02634), which is
also the Mooncake simulator's infinite-capacity ceiling; bounded
capacities give that simulator's [capacity
curve](https://kvcache.ai/blog/calculate-kvcache-cache-budge/). LRU is the
default eviction policy in both vLLM (`FreeKVCacheBlockQueue`) and SGLang
(`--radix-eviction-policy lru`), though SGLang evicts radix-tree leaves
rather than flat blocks.

Prefixes are identified by chained per-block hashing (`hash(parent,
block)`, 16-token blocks; 16 is the smallest block vLLM's FlashAttention
backend supports and the granularity it aligns to), following the
[qwen-bailian trace](https://github.com/alibaba-edu/qwen-bailian-usagetraces-anon)
convention.

### What does the curve assume?

One global cache, i.e. a single worker or perfect routing. With N workers
behind hash routing each sees a partition, so read it as fleet-level
guidance rather than per-worker sizing. An engine's KV pool also holds
in-flight requests; only the remainder is available for reuse.

The ceiling covers sharing between requests only. Intra-session reuse is
not measurable without session identifiers in the log.

### How do I convert token counts to gigabytes?

Per-token cost follows the architecture. GLM-5.2 caches two things:

| component | derivation | per token |
|---|---|---|
| MLA latent | `(512 + 64) x 2 bytes x 78 layers` | 87.75 KiB |
| DSA indexer | `128 x 2 bytes x 21 full-indexer layers` | 5.25 KiB |
| | | **93 KiB** |

Only 21 of the 78 layers run a full indexer, the rest reuse the previous
one's selection. Both terms match the `kvcache-simulator` catalog, which
reports 95,232 bytes per token for this model.

Do not scale that figure to another model. Only some architectures carry a
DSA indexer, and the attention term differs as well: Kimi K2.5 is 68.6 KiB
per token, a 32-layer fp16 GQA model 128 KiB. Look yours up with
`kvcache-simulator list-models`.

### Has this been checked against other tools?

Yes, twice, both on the public Mooncake conversation trace with 512-token
blocks, and both pinned by tests in `tests/test_core.py`.

**NVIDIA AIPerf v0.11.0.** `aiperf analyze-trace` reports
`cache_hit_rate: 0.38425808746366197`, which is the unweighted mean of
per-request hit-block ratios. Computing that same definition from our own
block matching yields `0.3842580874636671` — a difference of 5e-15, so the
two implementations agree on every request's prefix match.

**kvcache.ai `kvcache-simulator`.** It reports a 39.8% ceiling on this
trace, which our default now reproduces exactly, and an LRU curve of 4.5%
at 30 blocks, 5.5% at 1,937, 17.9% at 7,729 and 34.6% at 30,918 — the same
shape as ours at comparable block counts. Feeding it our own `export`
output works too: it reads the file and lands within 0.1 points of our
figure on the same data.

### Why measure over only the last half of the log?

Because the first half is the cache filling up. Counting it reports what
one cold replay of that particular log would have achieved, not what a
long-running service settles at. The simulator makes the same choice and
we follow it, so the two are directly comparable: both report 39.8% on the
Mooncake trace. `--warmup 0` gives the cold-start reading, 37.4% on the
same trace.

The AIPerf comparison is a separate axis. It averages per-request hit
ratios, weighting a 900-token request the same as a 126,000-token one,
while we divide total hit tokens by total input tokens. Token weighting is
the right choice for sizing a cache, since it is tokens that occupy it.

### How exact are the token counts?

Characters divided by four, unless the log carries `usage`, in which case
exact counts are used and the report says so. CJK-heavy text tokenizes
closer to 1-2 characters per token, so absolute counts run low while
ratios remain usable.

### What has not been verified?

- **Replay round-trip.** We have not fed an export through AIPerf or
  SGLang and confirmed it replays. The formats match their documented
  schemas, which is not the same thing.
- **Exact tokenization.** Approximate mode hashes character chunks, so
  reported structure is a lower bound on what a real tokenizer would match:
  formatting differences that tokenize identically read as distinct
  prefixes here.
- **Landscape claims.** The statement below that no tool goes from a raw
  gateway log to a hashed trace comes from a survey of vendor
  documentation in July 2026. That is an absence of evidence, not a vendor
  denial.

### Why not just use AIPerf or the Mooncake simulator?

If you already have a hashed trace, use them. `kvcache-simulator` does
more than we do on that input: per-model byte accounting across dozens of
models, FIFO/LRU/Optimal policy comparison, a C++ replay core. AIPerf is a
full benchmarking harness that actually drives an endpoint.

They start from a line like
`{"timestamp": 0, "input_length": 6909, "hash_ids": [1234, 5678, ...]}`.
A gateway log is not that — it is prompts. Turning one into the other
means extracting the prefill text, chunking it, salting and chaining the
hashes, and doing it on a machine you control because that step is the one
that touches real prompts. That is the step this tool covers, and its
`export` output is what those tools consume.

Reading the raw log also preserves two things a hashed trace has already
discarded: whether requests carry session identifiers, which decides
whether cheap affinity routing would work at all, and which specific
prefixes carry the reusable volume. Neither analyzer reports those,
because by the time they see the data it is gone.

## License

Apache-2.0
