Validating a model with a few lines of code

OceanVal has a built-in 3-step approach that will let you validate models using popular gridded datasets quickly. It will take care of everything for you: it works out which model variable maps to which observational variable, and it will matchup and validate each variable. This can all be done with 2 lines of code.

This is all done using OceanVal's built-in data recipes. If you want to know about the recipes available, read here.

oceanval.create_recipes does that for you. It scans your model output, works out which model variable holds each observational variable, and writes a complete, runnable matchup script. Validating a global or Northwest European Shelf model is then two steps: generate the script, then run it.

Note

This approach involves downloading observational data over the internet. In cases where this data is temporarily not available, OceanVal will let you know.

Step 1: Generate the script for model validation

The first step is to tell OceanVal where your simulation data can be found. It will then look through the directories netCDF files and create a script that can be used to matchup and validate your simulation.

python
import oceanval

oceanval.create_recipes(
    simdir="/path/to/model/output",
    ndown=2,
    out="matchup.py",
    domain="global",
    start=2005,
    end=2014,
)

Only one file per output stream is read, not every file in the simulation, so this takes seconds rather than minutes even for a long run.

How does create_recipes identify variables?

From the long_name attribute in your NetCDF files, not the variable name. A variable called N3_n with a long name of nitrate nitrogen is recognised as nitrate; votemper described as temperature is recognised as temperature. You don't have to rename anything or write a mapping table. Where a model splits a quantity across several variables — chlorophyll across phytoplankton functional types, say — they are summed.

Note: variable identification is an automated process and is not 100% guaranteed to identify variables correctly, so double check the output variables.

Step 2: Check what was found

create_recipes prints what it identified, so you can check it against what you know your model writes before running anything:

output
Wrote matchup.py
  ammonium: N4_n
  chlorophyll: P1_Chl+P2_Chl+P3_Chl+P4_Chl
  nitrate: N3_n
  oxygen: O2_o
  phosphate: N1_p
  salinity: so
  silicate: N5_s
  temperature: thetao
  commented out (no model variable found): alkalinity, kd490, ph

A name joined with + is a sum — here the four ERSEM phytoplankton chlorophyll variables, which together are what the satellite observes.

The last line is what OceanVal could not find. This run has no carbonate chemistry or optics, so alkalinity, pH and KD490 were left out. Nothing is silently dropped: those recipes are still written into the script, commented out and ready for you to complete. See Filling in the gaps.

Step 3: Run it

The generated file is an ordinary Python script. Read it, adjust anything you want to change, and run it:

bash
python matchup.py

The script registers every recipe it found a variable for, then calls matchup and validate for you. OceanVal will report the file pattern it has identified and ask you to confirm before matching. When the build finishes, the report opens in your browser.

Observations download themselves

The recipes fetch their observational data when matchup runs — you don't need to download WOA23, OCCCI, GLODAP or NSBC yourself. The first run will therefore spend time pulling data over the network.

Choosing the domain

Several variables have a recipe in both regions. Registering the same variable twice would replace the first registration, so only one recipe per variable can be live — domain decides which.

This is a default, not a restriction. A variable with a recipe in only one of the two regions still gets that recipe whichever domain you pick:

Variable Global NW European Shelf Affected by domain
TemperatureCOBE-SST 2, WOA23NSBCYes
SalinityWOA23NSBCYes
NitrateWOA23NSBCYes
PhosphateWOA23NSBCYes
SilicateWOA23NSBCYes
OxygenWOA23NSBCYes
ChlorophyllOCCCINSBCYes
Ammonium—NSBCNo
pHGLODAPv2.2016b—No
AlkalinityGLODAPv2.2016b—No
KD490OCCCI—No

So a shelf model run with domain="nwes" still gets pH, alkalinity and KD490 from the global datasets, and a global model run with domain="global" still gets ammonium from NSBC. The recipe that loses out is written into the script as a commented alternative, so you can swap back by uncommenting it and commenting out the other.

Inside the generated script

Every recipe OceanVal ships with appears in the script, in one of three states. The first is live — a model variable was found, and this is the recipe for your domain:

python
# Chlorophyll - North Sea Biogeochemical Climatology
# recipe: 'nsbc'
oceanval.add_gridded_comparison(
    name="chlorophyll",
    model_variable="P1_Chl+P2_Chl+P3_Chl+P4_Chl",
    recipe={"chlorophyll": "nsbc"},
    climatology=True,
    vertical=False,  # set True to validate the full water column
)

The second is an alternative source for a variable that is already covered, commented out with an explanation of what to do if you want it instead:

python
# Chlorophyll - Ocean Colour CCI (https://esa-oceancolour-cci.org/)
# recipe: 'occci'
# The Northwest European Shelf recipe for chlorophyll is registered instead.
# Registering a second one would replace it, so to validate against
# this source, comment out the other recipe and uncomment this one.
# oceanval.add_gridded_comparison(
#     name="chlorophyll",
#     model_variable="P1_Chl+P2_Chl+P3_Chl+P4_Chl",
#     recipe={"chlorophyll": "occci"},
#     climatology=False,
# )

The third is a recipe with no model variable found — see below. The script ends with the matchup and validate calls that finish the job, with sim_dir, n_dirs_down and the simulation's year range already filled in.

Filling in the gaps

A recipe OceanVal could not find a model variable for is written out commented, with the model variable left as a placeholder:

python
# Alkalinity - GLODAPv2.2016b (https://www.glodap.info/)
# recipe: 'glodap'
# No alkalinity variable was found in the model output.
# Set model_variable to the name your model uses, then uncomment
# this block.
# oceanval.add_gridded_comparison(
#     name="alkalinity",
#     model_variable="talk",
#     recipe={"alkalinity": "glodap"},
#     climatology=True,
# )

To use it, uncomment the block and replace model_variable with the name your model uses. There are two common reasons a variable is missed:

Before you run it

Two things in the generated script are worth a look before you run it.

WOA23 decadal periods. The World Ocean Atlas publishes temperature and salinity per decade, and start/end must sit inside one period. The script fills in the period covering your simulation; if your run straddles two, it says so and you will need to choose one.

Water-column validation. Recipes that support it are written with vertical=False, which validates the surface only. Set it to True to validate the full water column — and then matchup also needs thickness, either "z_level" or the name of your model's cell thickness variable.