# The live prompt-parity gate (card t_6de5fc53) run on the Tiel file — t_7c926398 receipt
#
# command (from the worktree at the committed corrected instrument):
#   systemd-run --user --unit=t7c9-live-parity --property=MemoryMax=infinity --collect --wait --pipe \
#     bash -c 'cd /var/home/rybens/workspace/ggufone-wt-t7c9 \
#       && export GGUFONE_RUNTIME_DIR=/home/rybens/.local/share/ggufone/runtime/b11026-linux-x64-vulkan \
#       && export GGUFONE_BENCH_MODEL=/var/home/rybens/.hermes/models/Tiel-Coder-35B-A3B-UD-Q4_K_XL.gguf \
#       && uv run --frozen --offline --extra dev pytest -q --run-network -s -p no:cacheprovider \
#            tests/test_bench_live.py::test_the_bench_sends_the_same_prompt_as_the_serving_path'
#
# captured stdout of the unit:

Running as unit: t7c9-live-parity.service; invocation ID: 20d6bb01b2a04d96830c775ba8937db5
load_backend: loaded RPC backend from /home/rybens/.local/share/ggufone/runtime/b11026-linux-x64-vulkan/libggml-rpc.so
[dlssnr-layer] === VK_LAYER_NV_dlssnr loaded (VKLayer_DLSS5=(unset)) ===
[dlssnr-layer] [layer] vkCreateInstance -> 0xcf62a10
ggml_vulkan: Found 1 Vulkan devices:
ggml_vulkan: 0 = NVIDIA GeForce RTX 3060 Ti (NVIDIA) | uma: 0 | fp16: 1 | bf16: 1 | fp4: 0 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: NV_coopmat2
load_backend: loaded Vulkan backend from /home/rybens/.local/share/ggufone/runtime/b11026-linux-x64-vulkan/libggml-vulkan.so
load_backend: loaded CPU backend from /home/rybens/.local/share/ggufone/runtime/b11026-linux-x64-vulkan/libggml-cpu-haswell.so
[dlssnr-layer] [layer] vkCreateDevice -> 0x11b5d660 on NVIDIA GeForce RTX 3060 Ti (inert=1 enabled=0)
~llama_context:    Vulkan0 compute buffer size is 704.5996 MiB, matches expectation of 704.5996 MiB
~llama_context: Vulkan_Host compute buffer size is  21.4873 MiB, matches expectation of  21.4873 MiB
~llama_context:    Vulkan0 compute buffer size is 704.5996 MiB, matches expectation of 704.5996 MiB
~llama_context: Vulkan_Host compute buffer size is  21.4873 MiB, matches expectation of  21.4873 MiB

item c01: serving prefix 109 tokens, bench prefix 109 tokens, framing chat-template: qwen35moe / builtin
.
1 passed in 38.80s
          Finished with result: success
Main processes terminated with: code=exited, status=0/SUCCESS
               Service runtime: 39.749s
             CPU time consumed: 55.609s
                   Memory peak: 18.2G (swap: 26.4M)
