Round 43 C1 Benchmark  forward vs forward_batched (vectorised _refine_hcrit)
Device: cuda (NVIDIA GeForce RTX 4060 Laptop GPU)
Warmup=2, best-of-5

       n    K    serial_ms     batch_ms    speedup
----------------------------------------------------
   50000    1       238.13       116.65       2.04x
   50000    4       979.64       318.31       3.08x
   50000    8      2035.18       550.50       3.70x
   50000   16      4050.14      1037.98       3.90x
   50000   32      7676.02      2268.07       3.38x
   50000   64     15506.93      4283.20       3.62x
  100000    1       198.96       115.24       1.73x
  100000    4       761.10       305.93       2.49x
  100000    8      1477.92       553.36       2.67x
  100000   16      2940.62      1048.92       2.80x
  100000   32      5840.39      2175.66       2.68x
  100000   64     11846.30      4287.32       2.76x
