2000 tissues, 500 repetitions, cpu

What resolving a structure costs, and what a bound call saves
                        sequence  resolve (s)  bound (ms)  unresolved (ms)
   500-repetition fingerprinting         1.94        1.18           346.30
              500-echo refocused         0.16        1.88           107.74

  A dictionary sweep, a fit or a design loop pays the resolve once and
  the bound column on every call after it. A single curve pays the
  resolve for nothing.

Fixed per-event cost against per-state cost
 orders   best (ms)   ms per order
      8        23.2           2.91
     16        25.8           1.61
     32        26.9           0.84
     64        38.2           0.60

  fixed 21.1 ms, 0.27 ms per order: the fixed part is 71% of a 32-order run.
  That is the whole of what hoisting the relaxation factors can save.

The real subspace against the complex path, same event stream
  in phase, real               47.9 ms
  quarter turn, complex       235.2 ms

Is the fast path reached? Auto against a forced verdict
        pass       auto  forced real  forced complex
     forward      20.5 ms      18.6 ms        226.9 ms
forward-mode      29.1 ms      20.6 ms        288.6 ms

  the real path is 4.9x the complex one on this machine.

What the Jacobian costs on each path, same event stream
  in phase, real           forward     35.2 ms   jacobian    207.7 ms     5.9x   gradient    163.4 ms     4.6x
  quarter turn, complex    forward    286.2 ms   jacobian   1702.2 ms     5.9x   gradient   2139.6 ms     7.5x

  A pass that carries a derivative beside every quantity should cost a
  few times the plain one. The multiple is the check that the dual and
  adjoint kernels take the same path the plain ones take, rather than
  the arithmetic being cheap where the plumbing is not. Run the whole
  file again under BLOCHSIM_REAL_SCALAR=1 to separate what the real
  subspace is worth from what the lane kernels on top of it are worth:
  a multiple that does not move is a pass with no laned kernel to take.

Again on the card, at a size that fills it:

100000 tissues, 500 repetitions, cuda

What resolving a structure costs, and what a bound call saves
                        sequence  resolve (s)  bound (ms)  unresolved (ms)
   500-repetition fingerprinting         4.69        4.23          1645.83
              500-echo refocused         0.55        4.68           632.24

  A dictionary sweep, a fit or a design loop pays the resolve once and
  the bound column on every call after it. A single curve pays the
  resolve for nothing.

Fixed per-event cost against per-state cost
 orders   best (ms)   ms per order
      8        66.8           8.35
     16       110.6           6.91
     32       199.7           6.24
     64       364.9           5.70

  fixed 24.2 ms, 5.32 ms per order: the fixed part is 12% of a 32-order run.
  That is the whole of what hoisting the relaxation factors can save.

The real subspace against the complex path, same event stream
  in phase, real              171.3 ms
  quarter turn, complex       378.0 ms

Is the fast path reached? Auto against a forced verdict
        pass       auto  forced real  forced complex
     forward     107.2 ms     106.3 ms        284.3 ms
forward-mode     193.8 ms     192.5 ms        496.0 ms

  the real path is 2.2x the complex one on this machine.

What the Jacobian costs on each path, same event stream
  in phase, real           forward    164.8 ms   jacobian   1069.5 ms     6.5x   gradient   1378.4 ms     8.4x
  quarter turn, complex    forward    385.9 ms   jacobian   2309.6 ms     6.0x   gradient   3017.5 ms     7.8x

  A pass that carries a derivative beside every quantity should cost a
  few times the plain one. The multiple is the check that the dual and
  adjoint kernels take the same path the plain ones take, rather than
  the arithmetic being cheap where the plumbing is not. Run the whole
  file again under BLOCHSIM_REAL_SCALAR=1 to separate what the real
  subspace is worth from what the lane kernels on top of it are worth:
  a multiple that does not move is a pass with no laned kernel to take.
