Skip to content
pmlmodellingPublic

About

Automated ocean model validation in Python

Resources

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

OceanVal

Documentation Status Conda Badge GitHub Testing


Ocean model validation made easy in Python.

To learn more about the package, visit the OceanVal website. The

OceanVal is designed for the automated creation of validation reports. You provide the model and validation data. OceanVal does the rest. A short example of what the report looks like can be found here.

Running OceanVal from the terminal

The browser window is in the development version only (main), not yet in a tagged release; there is a page and demo video on the website.

The quickest way needs no Python at all. In a terminal, in the directory to work in, run oceanval (or OceanVal). A page opens in your web browser that takes you through these steps:

  1. Choose to match up new data and validate it, to match up only, or to validate matchups made earlier. Add your own validation data for future use saves observations of your own as recipes in a .oceanvalrc file (see below), instead of running anything.
  2. Simulation: where the model output is, how many directories down its files are, which files to skip or keep (the file filters), and the domain. Type the directory, with folders suggested as you go, or pick it with Browse…, which looks through the folders on the machine OceanVal runs on, so it works on a remote machine too. As you type, the page counts the files that pass the filters. The next page starts with years inferred from their names, which you can change there. It also asks where to save the matchups and the report: the directory you started oceanval in, unless you choose another. To validate matchups made earlier, you choose the report's options instead.
  3. Own data: whether you have observations of your own, as well as OceanVal's built-in datasets. If you do, a page for point data (csv files) and one for gridded data (netCDF) let you add them one at a time, with every argument of add_point_comparison or add_gridded_comparison; the ones that have to be given are marked in red. Skip a page if you have none of that kind. The calls are written into the matchup script. Add this data for use now only adds a dataset to this run, and Add this data for use now and in future also saves it as a recipe (see below), after a page asking for what a recipe needs that the form did not, such as the units, with what OceanVal will assume in amber.
  4. Recipes: the recipes window, to check the model variables found and choose the observations to validate against. A dataset validated through the water column (Vertical) needs a thickness before you can carry on.
  5. Units: a table of the units of every gridded and point matchup, read from the model's netCDF files and, for the observations, from the recipes (the Units column of the recipes page) or from your own netCDF file. Where OceanVal thinks they differ, it suggests a multiplier and an adder for the observations, in red and bold, and says what it assumed, such as a seawater density of 1025 kg/m³ between per kilogram and per volume. It can get this wrong, so check every row, and fill in any it could not work out. Your own point data is csv files, which have no units, so it is up to you to make sure they match the model's. The units always have to be confirmed, with a box under the tables, before the button starts the matchup. The conversions are written into the matchup script as obs_multiplier and obs_adder.
  6. Files: the page says it is identifying the files that meet your criteria, then shows what it found as a table: each variable, its model variable, the observations it is compared with, the files it is in, with List all files for each (OceanVal applies temporal subsetting to them), and their time resolution (monthly, or 1d, 5d and so on; the finest is used). Carry on if the matchups are right; No stops the run; Back stops it too, before anything is matched up, and goes back to the units. If a point dataset is matched by day (its point_time_res includes day, as the default does) against output coarser than daily, a page then says most of its observations may go unmatched, and offers to change point_time_res for all such datasets (OceanVal suggests year and month) or for each one, or to keep it. The choice is written into the matchup script. At a terminal, matchup() asks the same.
  7. Report (to match up and validate): "One last thing... How would you like your validation report?" The report's options, asked for before anything is matched up: a subregion to validate, regional summaries, a transect to validate the gridded datasets along (a start and an end longitude and latitude, which must run north-south or east-west, and the page does not let you carry on otherwise), fixed colour scales, PDF and Word versions, a zipped copy, a concise or detailed report, and, if a point dataset is matched up with Vertical, the depth bins its depth summaries use (each can be changed or removed, and more added). The report is built beside the matchups, in the directory chosen in step 2. Back shows the matchups again, and keeps the options. The options are written into the matchup script's validate() call.
  8. Run: matchup, then, to validate as well, validate() with the report's options. The page says what the run is doing, and shows the output (which appears as it comes in the terminal) as the matchups are made and if the run fails; anything else OceanVal asks, such as whether to try again for observations a server could not supply, is asked in the page. With jupyter-book 2 or later, the matchup builds an interim validation report as it goes: the page links to it once the first matchup's page is in it. Each matchup's page is added as soon as the matchup is made, and the full report (with PDF and Word, if asked for) is built once they all are. The finished page links to the full report.

Every step after the first has Back, which keeps what you have entered: get as far as the report's options, spot a mistake, go back as far as the simulation to fix it, and carry on with everything else as you left it. Whatever you change replaces what was there, in the steps after it too. The recipes window comes back as you left it; go back before it, and the simulation is read again, so the window keeps your global settings and the rows you changed whose model variables are still in the output, ticks the datasets afresh if you changed the domain, and says in red and bold what it could not keep. The units always have to be confirmed again. Once anything is matched up, there is no going back.

Simulations you have validated before (remembered as soon as you confirm the matchups) are listed by the arrow at the right of the simulation directory box. They are remembered in a .oceanvalcache JSON file in your home directory (OCEANVALCACHE names another file), which OceanVal only adds to and a new version keeps.

Closing the browser window quits oceanval in the terminal, within about 15 seconds, stopping anything it is running, whichever browser you use. So do Quit (which is instant) and Ctrl+C.

Outside the window, matchup(live_validation=...) builds the same interim report (HTML only, in oceanval_interim_report), with True or a dict of validate()'s report options, before validate() builds the full report:

oceanval.matchup(sim_dir="/path/to/output", start=2000, end=2010,
                 live_validation={"lon_lim": [-20, 10], "lat_lim": [40, 65]})
oceanval.validate(lon_lim=[-20, 10], lat_lim=[40, 65], pdf=True)

To see how it works first, choose Try a demo, below the four choices on the first page. You choose one of three CMIP6 climate models, each from a different ESGF server (NorESM2-LM, MPI-ESM1-2-LR and UKESM1-0-LL), and one to three of its variables (sea surface temperature, surface nitrate and sea surface salinity) in a grid; only those files are downloaded (up to 50 MB, 26 MB and 460 MB for the three models). It then takes you through the same steps, after a short page of instructions (which asks whether you want the validation concise and fast, the default, or detailed but slower) and without the step for your own data (only OceanVal's own datasets are used), with the options filled in, in red and bold, including the model's name as the only file names to use, matching 2010 up with COBE-SST 2 (temperature) and the World Ocean Atlas 2023 (nitrate and salinity). The units step suggests converting the observed nitrate, which you check as you would for your own model. It shows how OceanVal works, not how to validate a climate model, which needs many years of output rather than one. It creates a directory called oceanval_demo, with the data, matchups and report in it: remove it when you have finished.

The script it runs is written first, to matchup.py unless you choose another name, so you can run it again later with python matchup.py. On a remote machine, open the link oceanval prints: VS Code forwards its port for you. Over plain SSH (from a Windows terminal, say), oceanval prints the ssh -L command to run in a new terminal on your own computer, then you open the link in your browser. It uses port 8765 if that is free, so the command is the same each time; --port chooses another.

Your own recipes

Observations you use again can be saved as recipes of your own, in a .oceanvalrc file: in the directory you run OceanVal from, for that directory, or in your home directory, for everywhere (the one in the directory wins). Choose Add your own validation data for future use in the oceanval window, then point (csv files) or gridded data (netCDF on this machine, on a THREDDS server or at a web address), or Add this data for use now and in future in a matchup's own data step. OceanVal opens the data to check it, and checks the recipe against its own and your others, before it is saved. A source called MySat for chlorophyll is then used like any other recipe, and is offered in the recipes window, tagged Yours:

oceanval.add_gridded_comparison(
    model_variable="chl",
    recipe={"chlorophyll": "mysat"},
)

See the recipes page for the file's format. Recipes of your own are in the development version only (main), not yet in a tagged release.

Using built-in recipes

You can register a standard observation climatology without manually listing all metadata:

import oceanval

oceanval.add_gridded_comparison(
    name="temperature",
    source="WOA23",
    model_variable="temp",
    recipe={"temperature": "woa23"},
    start=2005,
    end=2014,
    climatology=True,
)

This uses the v0.2.0 recipe system for datasets such as WOA23, NSBC, OCCCI and GLODAP.

There are also point (in-situ) recipes — currently on main, not yet in a tagged release — from the ICES Oceanographic database covering temperature, salinity, alkalinity, ammonium, chlorophyll, nitrate, oxygen, pH, phosphate and silicate. Rather than reading csv files, they download the observations during matchup, for the years being matched and the lon_lim/lat_lim area (or the model's own extent); see the recipes page for the full list with units:

oceanval.add_point_comparison(
    model_variable="votemper",
    recipe={"temperature": "ices"},
    vertical=True,
)

Point matchups made through the water column are summarised by depth in the report, by default in 0-10, 10-30, 30-60, 60-100, 100-150, 150-300, 300-600, 600-1000 and >1000 m. validate() takes your own as [min, max] pairs in metres, with a max of None for everything below the deepest:

oceanval.validate(depth_bins=[[0, 20], [20, 200], [200, None]])

Generating a matchup script

Deprecated: create_recipes will be removed in a future release, and warns when called. Run the oceanval command instead, which opens the same window and goes on to run the matchup and the report.

create_recipes writes the script for you. It scans your model output, works out which model variable holds each observational variable — matching on the netCDF long_name attributes, so your variables need not be named after the observations — and writes out every built-in gridded recipe. Recipes it found a model variable for are live, with that variable filled in; any it could not identify are commented out for you to complete by hand, unless you fill them in first in the window described below. If your output looks like raw FVCOM output (an unstructured mesh), it first asks you to confirm that, then writes fvcom=True into the script's matchup() call — pass fvcom=True or fvcom=False to skip the question. Where a variable has a recipe in more than one region (e.g. temperature), domain picks which one is left live — "global" or "nwes" (Northwest European Shelf) — and a variable with a recipe only outside that domain still gets that one. domain="nwes" also adds the ICES point recipes (see above) — they have no global equivalent, so they're left out entirely for domain="global" rather than written out commented.

import oceanval

oceanval.create_recipes(
    simdir="/path/to/model/output",
    ndown=2,   # how many directories down the output files sit
    out="matchup.py",
    domain="global",
    start=2005,   # passed straight through to matchup()
    end=2014,
)

Before the script is written, a page opens in your web browser so you can check and change all of this. It has a row for each observational variable, holding the model variable identified for it (which you can change; a pop-out under the box lists every variable in the model output with its long_name; search it, and tick several to sum them as var1+var2+var3, with the ones selected listed at the top) and tick-boxes for the gridded and point datasets available for it. They start ticked as the script would otherwise be written; tick several to validate a variable against each of them, or use Clear all selections to untick everything and start from none. Each ticked dataset can be given years of its own, and the ones whose observations are resolved in depth (WOA23, NSBC and ICES) can be set to Vertical, to validate the full water column rather than the surface alone.

Pass gui=False to write the script straight away, e.g. in a batch job. create_recipes then asks in the terminal for the model variable of anything it could not identify (press Enter to skip, or pass ask=False). With validate=False, the script's validate() call is written commented out, to build the report once the matchups are made. exclude and require leave out output files by name, as matchup() does (e.g. exclude="5d"), and are passed on to the script's matchup() call.

Comparing multiple validation outputs

To compare validation reports from multiple simulations:

import oceanval

oceanval.compare(
    model_dict={
        "model_a": "/path/to/model_a",
        "model_b": "/path/to/model_b",
    },
    view=True,
    ask=True,
)

This recreates the v0.2.0 comparison workflow and writes the shared comparison report to oceanval_comparison/compare/_build/html/index.html.

Installation

OceanVal should be used with Python versions 3.10-3.13.

You can install the latest release OceanVal from conda-forge as follows:

conda install conda-forge::oceanval

You can install the development version of OceanVal from GitHub using the following steps.

First, clone this directory:

git clone https://github.com/pmlmodelling/oceanval.git

Then move to this directory.

cd oceanval

Second, set up a conda environment. If you want the envionment to called something other than oceanval, you can change the name in the oceanval.yml file.

conda env create -f oceanval.yml

Activate this environment.

conda activate oceanval

Now, sometimes R package installs go wrong in conda. Run the following command to ensure Rcpp is installed correctly.

Rscript -e "install.packages('Rcpp', repos = 'https://cloud.r-project.org/')"

Now, install the package.

pip install .

```sh
conda activate oceanval

Now, install the package.

pip install .

Alternatively, install the conda environment and package using the following commands:

    conda env create --name oceanval -f https://raw.githubusercontent.com/pmlmodelling/oceanval/main/oceanval.yml
    conda activate oceanval​
    pip install git+https://github.com/pmlmodelling/oceanval.git​

About

Automated ocean model validation in Python

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages