llama_batched_bench: n_kv_max = 34048, n_batch = 2048, n_ubatch = 512, flash_attn = 1, is_pp_shared = 1, is_tg_separate = 0, n_gpu_layers = 99, n_threads = 10, n_threads_batch = 10

|    PP |     TG |    B |   N_KV |   T_PP s | S_PP t/s |   T_TG s | S_TG t/s |      T s |    S t/s |
|-------|--------|------|--------|----------|----------|----------|----------|----------|----------|
|   128 |    128 |    1 |    256 |    0.035 |  3655.16 |    0.620 |   206.50 |    0.655 |   390.91 |
|   128 |    128 |    2 |    384 |    0.021 |  6015.89 |    0.675 |   379.46 |    0.696 |   551.79 |
|   128 |    128 |    4 |    640 |    0.021 |  6022.11 |    0.993 |   515.52 |    1.014 |   630.90 |
|   128 |    128 |    8 |   1152 |    0.022 |  5952.93 |    1.687 |   606.93 |    1.709 |   674.21 |
|   128 |    128 |   16 |   2176 |    0.021 |  6011.65 |    1.091 |  1877.24 |    1.112 |  1956.38 |
|   128 |    128 |   32 |   4224 |    0.021 |  5991.39 |    1.523 |  2689.59 |    1.544 |  2735.27 |
|   128 |    128 |   64 |   8320 |    0.021 |  5960.70 |    2.571 |  3186.05 |    2.593 |  3209.03 |
|   128 |    128 |  128 |  16512 |    0.022 |  5902.97 |    5.180 |  3163.17 |    5.201 |  3174.59 |
|   128 |    128 |  256 |  32896 |    0.022 |  5743.52 |   11.662 |  2809.88 |   11.684 |  2815.48 |
|   128 |    256 |    1 |    384 |    0.023 |  5559.17 |    1.238 |   206.73 |    1.261 |   304.43 |
|   128 |    256 |    2 |    640 |    0.021 |  5992.51 |    1.355 |   377.83 |    1.376 |   464.96 |
|   128 |    256 |    4 |   1152 |    0.021 |  5974.05 |    2.005 |   510.73 |    2.026 |   568.50 |
|   128 |    256 |    8 |   2176 |    0.022 |  5916.61 |    3.418 |   599.24 |    3.439 |   632.69 |
|   128 |    256 |   16 |   4224 |    0.021 |  5967.09 |    2.225 |  1841.30 |    2.246 |  1880.71 |
|   128 |    256 |   32 |   8320 |    0.022 |  5939.68 |    3.208 |  2553.39 |    3.230 |  2575.99 |
|   128 |    256 |   64 |  16512 |    0.022 |  5881.00 |    5.785 |  2832.08 |    5.807 |  2843.51 |
|   128 |    256 |  128 |  32896 |    0.022 |  5783.74 |   12.871 |  2545.89 |   12.893 |  2551.45 |


0.00.494.912 W load: control-looking token: 128247 '</s>' was not control-type; this is probably a bug in the model. its type will be overridden
0.00.809.068 W llama_context: n_ctx_seq (34048) > n_ctx_train (32768) -- possible training context overflow