YAML-driven BSM and flavor scans

Define physics models declaratively. Scan them with reusable engines.

BSMScanner is a model-independent framework for Beyond-the-Standard-Model parameter scans. Models are written in YAML, validated and lowered in Python, evaluated through the framework pipeline, and saved as machine-readable artifacts for later analysis.

Quickstart

Install, test, and run

1. Install locally

python -m venv .venv
.venv/bin/python -m pip install -e .

2. Run tests

.venv/bin/python -m pytest tests -q

3. Run an example

.venv/bin/python examples/scotogenic_ma/run_scan.py \
  --model models/scotogenic_ma/model_no.yaml \
  --run-dir examples/scotogenic_ma/runs/demo

Framework Capabilities

What BSMScanner does

BSMScanner separates model physics from reusable framework logic. Model authors define parameters, functions, matrices, outputs, and likelihood terms. The framework handles validation, dependency ordering, diagonalization, scan orchestration, artifact writing, and statistics.

Declarative modelsYAML manifests with imports, parameters, observables, matrices, likelihoods, outputs, and scan settings.
Reusable core blocksShared neutrino, CKM, quark, constants, and likelihood-kernel infrastructure.
Flavor diagonalizationTakagi, SVD, Hermitian eigensolvers, PMNS, CKM, and optional flavor sectors.
Model-independent enginesRandom, SciPy DE, native adaptive DE, and basin orchestration scans.
Validity-aware outputsSeparate technical status from physics validity using status, valid, and failure_reason.
Plot-free statisticsMachine-readable likelihood-weighted CSV/JSON outputs ready for external plotting scripts.

Architecture

Core/model split

Core responsibilities

  • Reusable constants and observable blocks.
  • Generic matrix diagonalization infrastructure.
  • PMNS/CKM mixing-matrix construction.
  • Ordering-aware neutrino observable logic.
  • Generic likelihood kernels such as Gaussian, hard cut, and table terms.
  • Scan engine orchestration and artifact writing.

Model responsibilities

  • Scanned and fixed parameters.
  • Analytic functions, helper expressions, and mass matrices.
  • Matrix metadata: type, role, diagonalize.
  • Ordering choice and requested outputs.
  • Model-specific likelihood composition and numerical datasets.
  • Optional model-specific plugins or external backend calls.

Model Authoring

YAML model syntax

A model can be a single YAML file or a manifest importing reusable blocks. The preferred public pattern is a small manifest plus focused files for parameters, functions, matrices, likelihoods, outputs, and scan settings.

metadata:
  name: my_model
  version: 0.1.0

imports:
  - parameters.yaml
  - functions.yaml
  - matrices.yaml
  - ../../core/neutrino/normal.yaml
  - ../../core/neutrino/observables_common.yaml
  - likelihoods.yaml
  - outputs.yaml
  - scan.yaml

Parameters

parameters:
  - name: Retau
    value_type: real
    scan: true
    lower: -0.5
    upper: 0.5
    default: 0.0
    prior: flat

Mass matrices and diagonalization

matrices:
  Mnu:
    role: neutrino_mass
    type: complex_symmetric
    expression: [[...], [...], [...]]
    diagonalize:
      method: takagi
      output:
        masses: neutrino_masses
        unitary: U_nu

  Ml:
    role: charged_lepton_mass
    type: complex_general
    expression: [[...], [...], [...]]
    diagonalize:
      method: svd
      output:
        masses: charged_lepton_masses
        left_unitary: U_l_L
        right_unitary: U_l_R

Mixing matrices

mixing_matrices:
  PMNS:
    type: left_mismatch
    convention: U_left_dagger_U_right
    left: U_l_L
    right: U_nu
    output: U_PMNS

  CKM:
    type: left_mismatch
    convention: U_left_dagger_U_right
    left: U_u_L
    right: U_d_L
    output: V_CKM

Likelihood terms

likelihoods:
  - name: s13_sq_term
    kind: gaussian
    observable: s13_sq
    mean: 0.0222
    sigma: 0.0006

  - name: perturbativity_cut
    kind: hard_cut
    observable: max_coupling
    lower: 0.0
    upper: 12.566370614359172

Flavor System

Leptons, quarks, PMNS, and CKM

Lepton sector

Majorana neutrino matrices use Takagi factorization: U_nu^T Mnu U_nu = diag(m_i). Charged-lepton matrices use SVD-like biunitary diagonalization: U_L^† M U_R = diag(m_i).

PMNS is constructed as U_PMNS = U_l_L^† U_nu.

Quark sector

Up- and down-quark matrices are diagonalized with SVD. CKM is constructed as V_CKM = U_u_L^† U_d_L.

Reusable CKM observables include matrix-element absolute values, angles, CP phase, Jarlskog invariant, Wolfenstein parameters, and quark mass ratios.

Matrix Reconstruction

How diagonalization is declared and checked

BSMScanner stores diagonalization outputs as named quantities. A model declares the matrix, the method, and the names of the masses and rotation matrices. The same outputs can then be used for PMNS, CKM, observables, likelihoods, or external validation.

Majorana / Takagi

Use for complex symmetric neutrino mass matrices.

matrices:
  Mnu:
    role: neutrino_mass
    type: complex_symmetric
    expression:
      - [M11, M12, M13]
      - [M12, M22, M23]
      - [M13, M23, M33]
    diagonalize:
      method: takagi
      output:
        masses: neutrino_masses
        unitary: U_nu

Check: U_nu.T @ Mnu @ U_nu ≈ diag(neutrino_masses)

Dirac / SVD

Use for charged leptons and quarks with general complex matrices.

matrices:
  Ml:
    role: charged_lepton_mass
    type: complex_general
    expression:
      - [Ye11, Ye12, Ye13]
      - [Ye21, Ye22, Ye23]
      - [Ye31, Ye32, Ye33]
    diagonalize:
      method: svd
      output:
        masses: charged_lepton_masses
        left_unitary: U_l_L
        right_unitary: U_l_R

Check: U_l_L.conj().T @ Ml @ U_l_R ≈ diag(charged_lepton_masses)

Hermitian / eigh

Use for Hermitian matrices or derived objects such as M M†.

matrices:
  Hq:
    role: custom_mass_squared
    type: hermitian
    expression:
      - [H11, H12, H13]
      - [conj(H12), H22, H23]
      - [conj(H13), conj(H23), H33]
    diagonalize:
      method: hermitian_eigh
      output:
        eigenvalues: Hq_eigenvalues
        unitary: U_Hq

Check: U_Hq.conj().T @ Hq @ U_Hq ≈ diag(Hq_eigenvalues)

Constructing PMNS and CKM from diagonalization outputs

mixing_matrices:
  PMNS:
    type: left_mismatch
    convention: U_left_dagger_U_right
    left: U_l_L
    right: U_nu
    output: U_PMNS

  CKM:
    type: left_mismatch
    convention: U_left_dagger_U_right
    left: U_u_L
    right: U_d_L
    output: V_CKM

Python-side reconstruction check

During debugging, save the matrix and rotation outputs, then compare reconstruction numerically. Phases of unitary columns are conventional, so prefer reconstruction accuracy, masses, absolute mixing elements, and unitarity over raw element-by-element phase comparisons.

import numpy as np

def dagger(matrix):
    return np.asarray(matrix).conj().T

def check_takagi(Mnu, U_nu, masses):
    lhs = np.asarray(U_nu).T @ np.asarray(Mnu) @ np.asarray(U_nu)
    rhs = np.diag(masses)
    return np.max(np.abs(lhs - rhs))

def check_svd(M, U_L, U_R, masses):
    lhs = dagger(U_L) @ np.asarray(M) @ np.asarray(U_R)
    rhs = np.diag(masses)
    return np.max(np.abs(lhs - rhs))

def check_unitarity(U):
    U = np.asarray(U)
    return np.max(np.abs(dagger(U) @ U - np.eye(U.shape[0])))

def construct_pmns(U_l_L, U_nu):
    return dagger(U_l_L) @ np.asarray(U_nu)

def construct_ckm(U_u_L, U_d_L):
    return dagger(U_u_L) @ np.asarray(U_d_L)

Example output block for validation

outputs:
  save:
    - neutrino_masses
    - U_nu
    - charged_lepton_masses
    - U_l_L
    - U_l_R
    - U_PMNS
    - up_quark_masses
    - down_quark_masses
    - U_u_L
    - U_d_L
    - V_CKM
    - Vus
    - Vcb
    - theta12_q_deg
Compatibility fallback: neutrino-only models may use the documented charged-lepton identity fallback, so U_PMNS = U_nu. General lepton-sector models should define Ml and construct U_PMNS = U_l_L† U_nu.

Scan Engines

Engines, inputs, and basin orchestration

All engines receive the same compiled model evaluator, deterministic scanned-parameter order, parameter bounds and priors, objective mode, invalid-point policy, seed, selected outputs, and likelihood-term names. They differ only in how they propose parameter vectors.

EngineRule and techniqueInputs and free parametersOutputs
serial_random Native baseline engine. Draws independent bounded random points using linear, log, or signed-log prior coordinates. Requires max_evaluations > 0. Uses seed, objective, invalid_penalty, save_invalid_points, and save_every. Standard points.csv, metadata.json, best_fit.json, and summary.json.
diver External Diver bridge. BSMScanner supplies a bounded objective callback; Diver drives the population search. Requires a Diver-enabled build and max_generations or max_evaluations. Uses NP, convthresh, convsteps, and initialization controls. Standard BSMScanner artifacts. Diver raw file output is disabled by the bridge.
de_scipy Python reference backend using scipy.optimize.differential_evolution. Uses SciPy strategy, maxiter, popsize, tol, atol, mutation, recombination, init, polish, and x0. workers must be 1. Standard artifacts plus history.json with DE progress and convergence metadata.
adaptive_diver Native JADE-like current-to-pbest Differential Evolution with adaptive F/CR, archive support, bound handling, convergence checks, and optional SciPy local refinement. Requires population size at least four. Controls include population_size, max_generations, p_best_fraction, mutation/crossover ranges, bounds.handling, convergence, local refinement, and statistics. Standard artifacts plus history.json, final_population.csv, elite_points.csv, and optional population summaries.
basin_scan Model-agnostic orchestration: broad exploration, proposal/staged evaluation, selection, clustering, focused-box construction, optional refocus stages, then focused adaptive_diver runs. Uses exploration, selection, clustering, boxes, progressive exploration, proposals, refinement, manifold/ML focus, and focused-engine options. Focused engine is currently adaptive_diver. Standard artifacts plus exploration, selection, cluster, box, basin-ranking, progress, refocus, and focused-run artifacts.
Full manuscript: a complete academic reference for every engine and every basin_scan option is available at docs/scan_engines_manuscript.md.

de_scipy example

scan:
  engine: de_scipy
  settings:
    objective: nll
    invalid_penalty: 1.0e12
  de_scipy:
    strategy: rand1bin
    maxiter: 100
    popsize: 15
    tol: 0.01
    mutation: [0.5, 1.0]
    recombination: 0.7
    seed: 12345
    polish: false
    workers: 1

adaptive_diver example

scan:
  engine: adaptive_diver
  adaptive_diver:
    seed: 12345
    population_size: 80
    max_generations: 1500
    p_best_fraction: 0.1
    local_refinement:
      enabled: true
      method: Powell
      n_elites: 5
    output:
      save_history: true
      save_population: true
      save_elites: true

basin_scan with likelihood-balanced selection

scan:
  engine: basin_scan
  basin_scan:
    progressive_exploration:
      enabled: true
      n_rounds: 3
      points_per_round: [100000, 60000, 60000]
      selection:
        mode: balanced_terms
        total_top_fraction: 0.10
        term_quantile_cut: 0.30
        max_points: 2000
        min_points: 100
        terms: auto
        fallback_mode: top_fraction

    selection:
      mode: balanced_terms
      total_top_fraction: 0.10
      term_quantile_cut: 0.30
      max_points: 2000
      terms: auto

    focused_engine:
      name: adaptive_diver
      population_size: 100
      max_generations: 2500

basin_scan option catalogue

BlockRulePrincipal optionsMain outputs
exploration Draws the initial cloud in the full parameter box using Latin hypercube, Sobol, or uniform random sampling. method, n_points, keep_fraction. exploration_points.csv.
selection Keeps promising points by total objective, chi-square window, or likelihood-term-balanced cuts. mode, top_fraction, total_top_fraction, term_quantile_cut, terms, max_points, min_points, fallback_mode, near_miss. selected_points.csv, accepted_points.csv, near_miss_points.csv, selection_summary.json.
clustering Clusters selected points in normalized parameter coordinates with DBSCAN; falls back to one cluster if none is found. enabled, method: dbscan, eps_fraction, min_samples, max_clusters. clusters.csv.
boxes Builds quantile boxes around clusters, pads them, enforces minimum width, optionally clips and merges them. construction: quantile, q_low, q_high, padding_fraction, min_width_fraction, clip_to_original_bounds, merge_overlapping, max_boxes. focused_boxes.json.
progressive_exploration Replaces one-shot exploration with staged rounds that repeatedly select, box, and resample promising regions. enabled, n_rounds, points_per_round, nested selection, sampling.allocate_points, elite preservation, elite boxes, best-centered box, and round output switches. progressive_exploration_summary.json and round files under progressive_exploration/.
proposals / guided_sampling Transforms sampled points before evaluation and clips them to bounds. Stage type: prior_profile, complex_vector_norm, parameter_rescale, or point_function; plus probability, enabled, and apply_to. proposal_summary.json; proposal labels in point tables.
staged_evaluation Evaluates a cheap filtered graph first and fully evaluates only points passing the cheap-objective and hard-failure policy. Cheap include/exclude terms, theory checks, outputs; expensive terms; max_cheap_objective; require_no_hard_failures. staged_evaluation_summary.json and staged diagnostic columns.
refinement Jitters selected seeds, optionally reapplies proposals, evaluates new points, and reruns selection before clustering. enabled, n_rounds, points_per_seed, max_seeds, jitter_fraction, apply_proposals, keep_near_miss. refinement_points.csv, refinement_summary.json.
manifold_refocus Learns a covariance cloud in prior-aware transformed coordinates and converts sampled candidates into a smaller focused box. method: covariance, source, training caps, top_fraction_for_training, candidate count, covariance inflation, jitter, and focused-box quantiles. manifold_refocus_training.csv, manifold_refocus_candidates.csv, manifold_refocus_box.json, diagnostics.
ml_focus Trains an Extra Trees surrogate on evaluated points, ranks candidates, forms a focused box, and supplies optional seeds. model.type: extra_trees_regressor, estimator settings, training filters, candidate-source fractions, selected counts, focused-box options, and seed composition. ML focus diagnostics, candidates, selected points, focused box, and seeds.
focused_engine Runs one focused optimizer per focused box. name: adaptive_diver, population_size, max_generations, p_best_fraction, mutation/crossover, convergence, output, and local refinement options. basin_00/, basin_01/, ... and basin_results.json.

Statistics

Plot-free post-processing

BSMScanner intentionally does not generate plots inside scan engines. Engines produce complete machine-readable outputs; external scripts or notebooks handle visualization.

statistics:
  enabled: true
  method: de_weighted
  credible_levels: [0.68, 0.95]
  output_samples: true
  include_observables: true

de_weighted computes likelihood weights from evaluated DE points using exp[-0.5 * (chi2 - chi2_min)]. It is not a posterior sampler.

Artifacts

Standard output structure

run_directory/
  points.csv
  metadata.json
  best_fit.json
  summary.json
  history.json
  final_population.csv        # adaptive_diver when enabled
  elite_points.csv            # adaptive_diver when enabled
  exploration_points.csv      # basin_scan
  selected_points.csv         # basin_scan
  clusters.csv                # basin_scan
  focused_boxes.json          # basin_scan
  basin_results.json          # basin_scan
  selection_summary.json      # basin_scan
  statistics/
    de_weighted_samples.csv
    de_weighted_summary.json
    de_credible_intervals.json
    diagnostics.json

Important CSV columns include status, valid, failure_reason, scanner_target, metric_value, param::..., output::..., and generic likelihood components like__....

Examples

Runnable public examples

Scotogenic Ma

Radiative (one-loop) neutrino mass with a dark matter candidate.

.venv/bin/python examples/scotogenic_ma/run_scan.py \
  --model models/scotogenic_ma/model_no.yaml \
  --run-dir examples/scotogenic_ma/runs/demo

Minimal B-L

Gauged U(1)_B-L with a seesaw and a Z', the simplest single-file model.

.venv/bin/python examples/minimal_bl/run_scan.py \
  --run-dir examples/minimal_bl/runs/example_scan

SMEFT Wilson

SMEFT, Warsaw basis, 10 Wilson coefficients.

.venv/bin/python examples/smeft_wilson/run_scan.py \
  --run-dir examples/smeft_wilson/runs/demo

Z' Simplified

Z' simplified dark matter (LHC DM Forum benchmark).

.venv/bin/python examples/zprime_simplified/run_scan.py \
  --run-dir examples/zprime_simplified/runs/demo

Python API

Use BSMScanner from scripts

from pathlib import Path
from bsm_scanner import compile_model, load_model, run_scan

model = load_model(Path("models/minimal_bl/model.yaml"))
compiled = compile_model(model, build_backend=False)
results = run_scan(model, compiled, run_directory=Path("runs/minimal_bl"))

print(results.summary)
print(results.best_fit_path)

Reference Docs

More detailed files