date: 2026-09-15T09:39:23Z
os: Darwin 25.2.0 arm64
macos: 26.2 (25C56)
cpu: Apple M4 Pro
logical cpus: 14
performance cores: 10  efficiency cores: 4  smt: none
P-core L1d: 131072 B  L2 (shared per cluster of 5): 16777216 B
E-core L1d: 65536 B  L2 (shared per cluster of 4): 4194304 B
cache line: 128 B  page: 16384 B  memory: 24 GiB
isa features: CRC32 FlagM FlagM2 FHM DotProd SHA3 RDM LSE SHA256 SHA512 SHA1 AES PMULL SB FRINTTS PACIMP LRCPC LRCPC2 FCMA JSCVT PAuth PAuth2 FPAC FPACCOMBINE DPB DPB2 BF16 I8MM WFxT RPRES ECV AFP LSE2 CSV2 CSV3 DIT FP16 BTI SME SME2 SME_F64F64 SME_I16I64 
power: 'AC Power'
load average at start: 11.47 8.53 8.14
frequency: not published by the vendor for this part; see estimated clock below
compiler: Apple clang version 17.0.0 (clang-1700.4.4.1)
estimated clock: 4.50 GHz (dependent 1-cycle add chain, 400000000 adds, min of 7 runs)
bench: QUICK=0 REPS=default DOT_REPS=default CLOCK_GHZ=4.50
--- VECLIB_MAXIMUM_THREADS=1: every variant, one thread
sgemm: M = N = K = 1024 float32, 4194304 bytes per matrix, 2147483648 flops per call, reps 11 (plus 1 warmup, discarded)
blas: Accelerate cblas_sgemm, VECLIB_MAXIMUM_THREADS=1
RESULT sgemm_dim 1024 n
RESULT sgemm_matrix_bytes 4194304 bytes
RESULT sgemm_flops 2147483648 flops
RESULT sgemm_reps 11 reps
reference: naive C, checksum (sum of all entries) -8579.542461
RESULT sgemm_checksum -8579.542461 sum
naive_ijk                                min        2.534  median        2.545  max        2.559  cv   0.3%  n 11  GFLOP/s
RESULT naive_ijk_gflops 2.545 GFLOP/s
RESULT naive_ijk_gflops_max 2.559 GFLOP/s
RESULT naive_ijk_ms 843.696 ms
RESULT naive_ijk_cv 0.29 percent
RESULT naive_ijk_maxerr 0.000e+00 abs
ikj_autovec                              min       32.648  median       32.935  max       33.184  cv   0.4%  n 11  GFLOP/s
RESULT ikj_autovec_gflops 32.935 GFLOP/s
RESULT ikj_autovec_gflops_max 33.184 GFLOP/s
RESULT ikj_autovec_ms 65.203 ms
RESULT ikj_autovec_cv 0.45 percent
RESULT ikj_autovec_maxerr 0.000e+00 abs
blocked_neon_8x8                         min      110.570  median      111.766  max      112.600  cv   0.4%  n 11  GFLOP/s
RESULT blocked_neon_8x8_gflops 111.766 GFLOP/s
RESULT blocked_neon_8x8_gflops_max 112.600 GFLOP/s
RESULT blocked_neon_8x8_ms 19.214 ms
RESULT blocked_neon_8x8_cv 0.45 percent
RESULT blocked_neon_8x8_maxerr 0.000e+00 abs
blas_threads_1                           min     1518.551  median     1588.130  max     1605.196  cv   1.4%  n 11  GFLOP/s
RESULT blas_threads_1_gflops 1588.130 GFLOP/s
RESULT blas_threads_1_gflops_max 1605.196 GFLOP/s
RESULT blas_threads_1_ms 1.352 ms
RESULT blas_threads_1_cv 1.42 percent
RESULT blas_threads_1_maxerr 0.000e+00 abs
RESULT blas_threads_1_cpu_over_wall 1.00 ratio
fmadd_latency                            min        0.690  median        0.692  max        0.697  cv   0.4%  n 11  ns
RESULT fmadd_latency_ns 0.6897 ns
RESULT fmadd_latency_cv 0.36 percent
neon_fma_peak                            min      126.794  median      129.118  max      132.818  cv   1.3%  n 11  GFLOP/s
RESULT neon_fma_peak_gflops 129.118 GFLOP/s
RESULT neon_fma_peak_cv 1.31 percent
neon_fma_peak checksum 6.442451e+09
sgemm: every variant matched the naive result to within 1e-2
dot: 16777216 elements streaming (33554432 bytes int8 pair, 134217728 bytes float32 pair); 8192 elements L1-resident (16384 and 65536 bytes) times 2048 passes; reps 31 (plus 1 warmup)
dot references: int8 stream 4265427, int8 l1 -1948, f32 stream -1484.916734, f32 l1 13.870292
RESULT dot_int8_checksum 4265427 sum
RESULT dot_f32_checksum -1484.916734 sum
RESULT dot_stream_elements 16777216 elements
RESULT dot_stream_bytes_int8 33554432 bytes
RESULT dot_stream_bytes_f32 134217728 bytes
RESULT dot_l1_elements 8192 elements
RESULT dot_l1_bytes_int8 16384 bytes
RESULT dot_l1_bytes_f32 65536 bytes
RESULT dot_l1_passes 2048 passes
RESULT dot_reps 31 reps
dot_int8_sdot_stream                     min      116.189  median      118.602  max      124.989  cv   1.6%  n 31  ops/ns
RESULT dot_int8_sdot_stream_opsns 118.602 ops/ns
RESULT dot_int8_sdot_stream_cv 1.64 percent
dot_f32_fmla_stream                      min       29.190  median       29.879  max       30.648  cv   1.0%  n 31  ops/ns
RESULT dot_f32_fmla_stream_opsns 29.879 ops/ns
RESULT dot_f32_fmla_stream_cv 0.99 percent
RESULT dot_f32_fmla_stream_err_over_tol 0.001 fraction
dot_int8_sdot_l1                         min      185.043  median      191.012  max      204.238  cv   2.6%  n 31  ops/ns
RESULT dot_int8_sdot_l1_opsns 191.012 ops/ns
RESULT dot_int8_sdot_l1_cv 2.61 percent
dot_f32_fmla_l1                          min       46.722  median       47.513  max       49.008  cv   1.0%  n 31  ops/ns
RESULT dot_f32_fmla_l1_opsns 47.513 ops/ns
RESULT dot_f32_fmla_l1_cv 0.99 percent
RESULT dot_f32_fmla_l1_err_over_tol 0.000 fraction
dot: every result matched its reference
--- MODE=blas, VECLIB_MAXIMUM_THREADS unset: the vendor BLAS at its default thread count
sgemm: M = N = K = 1024 float32, 4194304 bytes per matrix, 2147483648 flops per call, reps 11 (plus 1 warmup, discarded)
blas: Accelerate cblas_sgemm, VECLIB_MAXIMUM_THREADS=(unset)
RESULT sgemm_dim 1024 n
RESULT sgemm_matrix_bytes 4194304 bytes
RESULT sgemm_flops 2147483648 flops
RESULT sgemm_reps 11 reps
reference: naive C, checksum (sum of all entries) -8579.542461
RESULT sgemm_checksum -8579.542461 sum
blas_threads_default                     min     3152.658  median     3180.082  max     3228.085  cv   0.7%  n 11  GFLOP/s
RESULT blas_threads_default_gflops 3180.082 GFLOP/s
RESULT blas_threads_default_gflops_max 3228.085 GFLOP/s
RESULT blas_threads_default_ms 0.675 ms
RESULT blas_threads_default_cv 0.74 percent
RESULT blas_threads_default_maxerr 0.000e+00 abs
RESULT blas_threads_default_cpu_over_wall 1.97 ratio
sgemm: every variant matched the naive result to within 1e-2
