{
  "model": "llama-3.1-8b",
  "fits": true,
  "arithmetic": "Llama 3.1 8B @ q8, 32,768 ctx, concurrency 1\n\n  weights   8.03B params x 8 bits / 8           =    7.5 GB\n  kv cache  131,072 B/tok x 32,768 tok x 1   =    4.0 GB   (GQA)\n  overhead  8% of weights + 1.5 GB floor      =    2.1 GB\n  --------------------------------------------------------------\n  required                                            13.6 GB\n  usable                                              96.0 GB\n  headroom                                            82.4 GB\n\n  decode is bandwidth-bound, and reads both:\n    weights    7.5 GB/token  (8.03B active)\n    kv cache   4.0 GB/token  (the whole cache, every token)\n    total     11.5 GB/token at 72% of peak\n  ~32 tok/s single-stream  (roofline estimate, not measured)"
}
