kv_cache_decode — LLMCORE decode op

• Data kinds: tokens × tokens × tokens → tokens

• Call: import fullseye as fs; fs.ledger.kv_cache_decode(query_step, key_cache, value_cache, scale=None) -> 'np.ndarray' (to call the implementation directly, import llmcore; llmcore.kv_cache_decode(query_step, key_cache, value_cache, scale=None) -> 'np.ndarray'; from the registry, opsllmcore.get("kv_cache_decode"))

Usage

> This operator's description has not been translated yet. The original text follows as it is.

1 トークン分だけ進める推論(KV Cache)→ tokens (1, dv)。

因果マスクの下では未来を書き換えても前の出力が 1 ビットも動かない(実測 0.0e+00)

ので、済んだ key/value を取っておいて 1 行だけ計算すればよい —— 逐次に 1 行ずつ

進めた結果は、全系列を一括で通した結果と一致する(相対差 1.6e-16)。

計算量は 1 歩あたり O(T d) で、一括の O(T² d) を T 歩に割ったものに等しい。

Args:

query_step: (1, d) float。いま生成しようとしている 1 トークン。

key_cache: (T, d) float。ここまでに見た key(自分を含む)。

value_cache: (T, dv) float。同じ長さの value。

scale: スコアの倍率。None なら 1/√d。

Returns:

tokens (1, dv) float: 一括で通した最終行と一致する。

Detailed usage guide

• llmcore family guide

References (sample data, literature)

• Sample-data catalog (download URLs / licences) — 2-D uses skimage.data (BSD/public domain) plus synthetic images; 3-D lists download URLs for real data sources (Stanford, PDS, …).

• Operator provenance and references — the sources of the research/methods this op family came from.

• The canonical algorithm (author, year) and its uses are named in the family usage guide above.

Runnable examples (verified samples that actually call this op)

• poc_attention_identities — py -3.11 examples/poc_attention_identities.py

Ops the type connects to (they accept tokens as input)

rms_norm · rope_rotate · attention_scores · attention_apply · attention_softmax · attention_tiled · attention_linear · attention_grouped

Same category (decode)

—


*Provenance: llmcore.py — LLMCORE operator registry. This per-op note is generated by tools/opdocs.py md (do not hand-edit).*

© 2026 Kazufumi Furuse — Fullseye operator documentation. Licensed under Apache-2.0.