bunker-stats · Notebook UX

Every helper in bunker_stats.notebook, run on one messy sample table and rendered exactly as it appears in a notebook. Numbers come from the Rust kernels; the layer only validates inputs and handles NaN/∞.

pip install "bunker-stats-rs[notebook]"

robust_summary

→ DataFrame
>>> nb.robust_summary(df)

Robust + classical descriptives per column: count, mean, std, median, MAD, IQR, Qn scale, trimmed mean, skew, kurtosis. Non-finite values are dropped and counted in n_missing.

  n n_missing mean std min median max mad mad_std iqr qn_scale trimmed_mean skew kurtosis
column                            
price 39.000 1.000 102.279 21.591 38.000 100.157 210.000 5.151 7.637 11.321 8.864 101.295 2.586 15.890
latency 40.000 0.000 49.168 5.256 36.175 49.015 57.856 3.578 5.305 6.964 5.526 49.408 -0.384 -0.203
demand 38.000 2.000 302.674 70.073 104.910 301.744 646.553 17.014 25.225 40.185 31.143 299.449 2.576 14.914
temperature 40.000 0.000 21.500 0.000 21.500 21.500 21.500 0.000 0.000 0.000 0.000 21.500 · ·

describe_fast

→ DataFrame
>>> nb.describe_fast(df)

A faster, richer df.describe().T backed by the Rust kernels. Adds a robust block (MAD, IQR, Qn, trimmed mean, skew, kurtosis) alongside the quartiles.

  n n_missing mean std min 25% 50% 75% max mad iqr qn_scale trimmed_mean skew kurtosis
column                              
price 39.000 1.000 102.279 21.591 38.000 95.421 100.157 106.742 210.000 5.151 11.321 8.864 101.295 2.586 15.890
latency 40.000 0.000 49.168 5.256 36.175 46.017 49.015 52.981 57.856 3.578 6.964 5.526 49.408 -0.384 -0.203
demand 38.000 2.000 302.674 70.073 104.910 277.826 301.744 318.012 646.553 17.014 40.185 31.143 299.449 2.576 14.914
temperature 40.000 0.000 21.500 0.000 21.500 21.500 21.500 21.500 21.500 0.000 0.000 0.000 21.500 · ·

outlier_report

→ DataFrame
>>> nb.outlier_report(df, method="iqr", k=1.5)

Per-column outlier counts, percentages and fence bounds. Methods: iqr, zscore, robust_zscore (median/MAD fences).

  method n n_missing n_outliers pct_outliers lower_bound upper_bound min max
column                  
price iqr 39.000 1.000 2.000 5.128 78.440 123.724 38.000 210.000
latency iqr 40.000 0.000 0.000 0.000 35.571 63.427 36.175 57.856
demand iqr 38.000 2.000 2.000 5.263 217.549 378.289 104.910 646.553
temperature iqr 40.000 0.000 0.000 0.000 21.500 21.500 21.500 21.500

normality_report

→ DataFrame
>>> nb.normality_report(df)

Jarque-Bera and Anderson-Darling diagnostics. The normal/conclusion verdict uses the JB p-value; A-D reports its statistic only (no p-value in the kernel).

  n skewness kurtosis jb_statistic jb_pvalue ad_statistic normal conclusion
column                
price 39.0000 2.5861 18.8896 453.7502 0.0000 5.1879 False reject normality (p < 0.05)
latency 40.0000 -0.3845 2.7967 1.0542 0.5903 0.2314 True cannot reject normality (p >= 0.05)
demand 38.0000 2.5755 17.9141 394.1945 0.0000 4.5595 False reject normality (p < 0.05)
temperature 40.0000 · · · · · · inconclusive (sample too small)

correlation_report

→ DataFrame
>>> nb.correlation_report(df, pvalues=True)

Correlation between numeric columns using pairwise-complete rows. Long form (pvalues=True) gives one row per pair with the test statistic and p-value; the default returns a square matrix.

  x y n correlation statistic pvalue
0 price latency 39.0000 -0.0917 -0.5604 0.5786
1 price demand 38.0000 0.9870 36.8608 0.0000
2 price temperature 39.0000 · · ·
3 latency demand 38.0000 -0.1542 -0.9365 0.3552
4 latency temperature 40.0000 · · ·
5 demand temperature 38.0000 · · ·

missingness_report

→ DataFrame
>>> nb.missingness_report(df)

The data-quality audit: separates NaN from ±inf from finite, per column, for every dtype. The one helper that counts rather than drops.

  dtype n_rows n_missing pct_missing n_finite n_infinite
column            
price float64 40.0 1.0 2.5 39.0 0.0
latency float64 40.0 0.0 0.0 40.0 0.0
demand float64 40.0 1.0 2.5 38.0 1.0
temperature float64 40.0 0.0 0.0 40.0 0.0
region str 40.0 0.0 0.0 · ·

rolling_report

→ DataFrame
>>> nb.rolling_report(df, "price", window=5)

Rolling-window features via the fused Rust kernel (all stats in one pass). Right-aligned and index-preserving. Showing the first 10 rows.

  price_roll5_mean price_roll5_std price_roll5_min price_roll5_max
0 · · · ·
1 · · · ·
2 · · · ·
3 · · · ·
4 99.690 8.058 85.677 105.169
5 124.554 47.803 100.080 210.000
6 127.086 46.500 103.575 210.000
7 125.318 47.769 94.735 210.000
8 124.464 48.254 94.735 210.000
9 123.706 48.685 94.735 210.000

bootstrap_ci_report

→ DataFrame
>>> nb.bootstrap_ci_report(df, stat="mean", random_state=0)

Bootstrap point estimate and confidence interval per column, via BootstrapConfig. Deterministic given random_state.

  stat n n_missing estimate ci_lower ci_upper conf
column              
price mean 39.000 1.000 102.289 96.170 109.932 0.950
latency mean 40.000 0.000 49.180 47.578 50.755 0.950
demand mean 38.000 2.000 302.767 282.269 327.801 0.950
temperature mean 40.000 0.000 21.500 21.500 21.500 0.950

scale_columns

→ DataFrame (+cols)
>>> nb.scale_columns(df, ["price", "latency"], method="robust")

Batch scaling (robust / zscore / minmax). Fit on finite values only, scattered back so NaN positions and row order are preserved. Showing the new columns beside their sources.

  price price_robust latency latency_robust
0 85.677 -1.896 57.206 1.544
1 100.080 -0.010 45.921 -0.583
2 103.575 0.447 49.897 0.166
3 105.169 0.656 49.339 0.061
4 103.949 0.496 42.707 -1.189
5 210.000 14.382 44.125 -0.922
6 112.736 1.647 51.067 0.387
7 94.735 -0.710 54.059 0.951
8 100.900 0.097 37.483 -2.174
9 100.157 0.000 56.109 1.337

winsorize_columns

→ DataFrame (+cols)
>>> nb.winsorize_columns(df, ["price"], lower_q=0.05, upper_q=0.95)

Batch winsorization — clip tails at the given quantile fractions. Note how the 210 and 38 outliers are pulled to the fences while other rows are untouched.

  price price_winsor
0 85.677 85.677
1 100.080 100.080
2 103.575 103.575
3 105.169 105.169
4 103.949 103.949
5 210.000 113.084
6 112.736 112.736
7 94.735 94.735
8 100.900 100.900
9 100.157 100.157
10 92.765 92.765
11 95.535 95.535
12 38.000 85.668
13 105.966 105.966

outlier_style

→ Styler
>>> nb.outlier_style(df, ["price", "demand"])

Highlights outlier cells across many numeric columns at once (red). Non-finite cells are left unstyled. Showing the first 16 rows.

  price latency demand temperature region
0 85.677161 57.205678 245.648614 21.500000 north
1 100.080222 45.921098 308.678150 21.500000 south
2 103.574811 49.897322 299.385059 21.500000 north
3 105.168638 49.338557 315.064689 21.500000 south
4 103.948634 42.707021 323.373571 21.500000 north
5 210.000000 44.125381 646.553473 21.500000 south
6 112.735866 51.067277 334.582161 21.500000 north
7 94.734514 54.059135 274.201344 21.500000 south
8 100.900079 37.483477 inf 21.500000 north
9 100.157292 56.108554 297.909445 21.500000 south
10 92.765327 56.982864 254.893382 21.500000 north
11 95.535345 52.835344 273.643931 21.500000 south
12 38.000000 50.166631 104.909720 21.500000 north
13 105.966229 50.867886 308.930736 21.500000 south
14 95.306842 36.175237 289.990367 21.500000 north
15 108.738903 48.062547 323.725582 21.500000 south

corr_heatmap

→ Styler
>>> nb.corr_heatmap(df)

Correlation matrix as a diverging background-gradient Styler (needs matplotlib). price↔demand shows the planted correlation.

  price latency demand temperature
column        
price 1.000000 -0.091734 0.987010 nan
latency -0.091734 1.000000 -0.154222 nan
demand 0.987010 -0.154222 1.000000 nan
temperature nan nan nan 1.000000

style_significance

→ Styler
>>> nb.style_significance(results, alpha=0.05)

Shades a results table by significance tier (p<α/50, p<α/5, p<α, then non-significant). NaN p-values stay unstyled.

  comparison pvalue cohens_d
0 A vs B 0.000300 1.350000
1 A vs C 0.011000 -0.720000
2 B vs C 0.048000 0.340000
3 C vs D 0.210000 0.110000
4 D vs E 0.830000 0.020000

style_effect_size

→ Styler
>>> nb.style_effect_size(results, "cohens_d")

Shades an effect-size column by |magnitude| against thresholds (default Cohen's 0.2/0.5/0.8: negligible→small→medium→large).

  comparison pvalue cohens_d
0 A vs B 0.000300 1.350000
1 A vs C 0.011000 -0.720000
2 B vs C 0.048000 0.340000
3 C vs D 0.210000 0.110000
4 D vs E 0.830000 0.020000

demean_style

→ Styler
>>> nb.demean_style(df, "price")

Legacy single-column styler (hardened): adds a demeaned column and colors each cell above (green) / below (red) the mean. Showing 14 rows.

  price latency demand temperature region price_demeaned
0 85.677161 57.205678 245.648614 21.500000 north -17.840276
1 100.080222 45.921098 308.678150 21.500000 south -3.437215
2 103.574811 49.897322 299.385059 21.500000 north 0.057374
3 105.168638 49.338557 315.064689 21.500000 south 1.651201
4 103.948634 42.707021 323.373571 21.500000 north 0.431197
5 210.000000 44.125381 646.553473 21.500000 south 106.482563
6 112.735866 51.067277 334.582161 21.500000 north 9.218429
7 94.734514 54.059135 274.201344 21.500000 south -8.782923
8 100.900079 37.483477 inf 21.500000 north -2.617358
9 100.157292 56.108554 297.909445 21.500000 south -3.360145
10 92.765327 56.982864 254.893382 21.500000 north -10.752110
11 95.535345 52.835344 273.643931 21.500000 south -7.982092
12 38.000000 50.166631 104.909720 21.500000 north -65.517437
13 105.966229 50.867886 308.930736 21.500000 south 2.448792

zscore_style

→ Styler
>>> nb.zscore_style(df, "price", threshold=2.0)

Adds a z-score column and highlights scores beyond ±threshold (high = orange, low = blue). Showing 14 rows.

  price latency demand temperature region price_zscore
0 85.677161 57.205678 245.648614 21.500000 north -0.503367
1 100.080222 45.921098 308.678150 21.500000 south -0.096982
2 103.574811 49.897322 299.385059 21.500000 north 0.001619
3 105.168638 49.338557 315.064689 21.500000 south 0.046589
4 103.948634 42.707021 323.373571 21.500000 north 0.012166
5 210.000000 44.125381 646.553473 21.500000 south 3.004427
6 112.735866 51.067277 334.582161 21.500000 north 0.260100
7 94.734514 54.059135 274.201344 21.500000 south -0.247812
8 100.900079 37.483477 inf 21.500000 north -0.073849
9 100.157292 56.108554 297.909445 21.500000 south -0.094807
10 92.765327 56.982864 254.893382 21.500000 north -0.303373
11 95.535345 52.835344 273.643931 21.500000 south -0.225216
12 38.000000 50.166631 104.909720 21.500000 north -1.848588
13 105.966229 50.867886 308.930736 21.500000 south 0.069093

iqr_outlier_style

→ Styler
>>> nb.iqr_outlier_style(df, "price", k=1.5)

Single-column IQR outlier highlight — a thin wrapper over the multi-column outlier_style. Showing 16 rows.

  price latency demand temperature region
0 85.677161 57.205678 245.648614 21.500000 north
1 100.080222 45.921098 308.678150 21.500000 south
2 103.574811 49.897322 299.385059 21.500000 north
3 105.168638 49.338557 315.064689 21.500000 south
4 103.948634 42.707021 323.373571 21.500000 north
5 210.000000 44.125381 646.553473 21.500000 south
6 112.735866 51.067277 334.582161 21.500000 north
7 94.734514 54.059135 274.201344 21.500000 south
8 100.900079 37.483477 inf 21.500000 north
9 100.157292 56.108554 297.909445 21.500000 south
10 92.765327 56.982864 254.893382 21.500000 north
11 95.535345 52.835344 273.643931 21.500000 south
12 38.000000 50.166631 104.909720 21.500000 north
13 105.966229 50.867886 308.930736 21.500000 south
14 95.306842 36.175237 289.990367 21.500000 north
15 108.738903 48.062547 323.725582 21.500000 south