Skip to content

Feature performance - #26

Merged
irm-codebase merged 37 commits into
mainfrom
feature-performance
Sep 24, 2026
Merged

irm-codebase merged 37 commits into
mainfrom
feature-performance

Conversation

@sjpfenninger

Copy link
Copy Markdown
Collaborator

Fixes #15

Summary of changes in this pull request

The main purpose of this PR was to address #15 and to optimise performance, but it bundles a couple of other changes too.

Claude Code was used for the memory and performance optimisations. In order to ensure these changes do not change results, a unit test suite checks against known-good data for a synthetic reference case, and for the Netherlands example already used in the integration test.

The following fixes and behaviour changes are also included:

  • protected is now a fraction of a pixel (0-1), but can still be treated as a binary layer to keep the previous behaviour 1:1.
  • The previous few PRs were not fully functional and are now fixed: scenario overrides were non-functional, ship-travel was looking for the non-scenario-aware top-level techs key, and the ship density downloads were failing unless forced to HTTP 1.1.

Performance improvements included:

  • Protected areas processing: now only reads those features that intersect the supplied reference raster, and uses streaming in batches via OGR+GDAL rather than loading everything at once.
  • Resampling: now reads inputs in windows and deals with arrays directly, with layers written to the NetCDF one at a time. To optimise memory use it uses float32 rather than float64. A land-cover lookup table is used instead of string replacement.
  • Area potential calculations: build a "keep mask" iteratively which is then applied in one go, and also switch to float32 output.
  • Reporting: strip-by-strip bincount aggregation reduces the need for keeping everything in memory at once, and the plotting is now memory-optimised and moved into separate rules as it proved to be particularly memory intensive.
  • Snakemake-specific improvements: threads: is passed through to GDAL; resources: mem_mb allows constraining parallelisation on memory-limited machines, and the memory-intensive rules also write a benchmark file for further monitoring.

Testing added:

  • New test-unit pixi environment pinned to the same versions as workflow/envs/module.*.pin.txt (sync enforced by a test)
  • Reference data regeneration via pixi run update-reference; diffs to tests/reference/ are intentional and should be reviewed.

Reviewer checklist

  • There are no pip dependencies in the module's environment files (workflow/envs/).
  • All rules use pathvars (e.g., <results>) in their inputs and outputs.
  • The integration test-suite is successful, including:
    • pre-commit.ci tests pass.
    • tests pass for all relevant OS configurations (linux, osx, windows).
  • Module documentation is up-to-date, including:
    • INTERFACE.yaml mentions all relevant pathvars and wildcards.
    • README.md describes how to use the module and has the necessary citations.

sjpfenninger and others added 30 commits September 8, 2026 16:00
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…path

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…etry buffer helper

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…eometries

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
sjpfenninger and others added 3 commits September 10, 2026 12:56
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

@irm-codebase irm-codebase left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good overall, but the unit tests need better environment isolation.

Details

In terms of memory efficiency this is night and day. Really impressive.
Some areas of the code seem a bit over optimised (e.g., plotting). It's ok to keep that as long as it does not become a maintenance burden.

General tests carried out:

  • Executed the integration test.
  • Executed for Mexico at state resolution (30+ shapes). Took ~13 minutes w/ 8 cores.

dask handles larger than memory operations well, as expected:

Image

Comment thread pixi.toml Outdated
Comment thread pixi.toml Outdated
Comment thread README.md Outdated
Comment thread pixi.toml Outdated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

consider renaming this to _plots.py or something similar, as everything here is related to plotting.

Comment on lines +90 to +93
# This call contains two optimisations for speed:
# 1. fig.savefig (rather than plt.savefig) disables pyplot's post-save re-render.
# 2. Omitting bbox_inches="tight" skips a measurement pre-render step.
fig.savefig(savefig, dpi=300)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I agree on using fig.savefig. Keeping bbox_inches="tight" might still be useful though, to avoid too much white space.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It seems that bbox_inches="tight" requires rendering the whole figure an additional time, so it does have quite an impact..

Comment thread tests/integration_test.py Outdated
Comment thread tests/integration_test.py
)
assert process.returncode == 0, process.stdout + process.stderr
return process.stdout + process.stderr

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tests from here on seem a bit over engineered. We can keep, but we should not be overly attached to them.

Comment thread unit_tests/conftest.py
@irm-codebase

Copy link
Copy Markdown
Contributor

This PR fixes hard-blocking issues related to this new feature.
#28

@jnnr jnnr left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for this! I am done with a first review. Test runs are running at the moment - I will check if the old results are reproduced for a European setting. I looked at most relevant parts, except for scripts/area_potential.py and scripts/resample.py. I will inspect them in a second round.

One of my central concern among the comments is that the test oracles probably do not give the validity that we desire for unit tests. In my view, it would be better to establish a source of truth can be easily checked by a human.

Comment thread unit_tests/conftest.py
Comment thread README.md Outdated
return config


def sequential_area_potential(ds, config):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This oracle repeats the code implementing the logic of get_area_potential


Therefore, limited validity is provided.

A better approach would be to define and document expected results for simple examples that can be easily verified visually. For example, see https://github.com/modelblocks-org/gregor/blob/main/test/_files/test.png

Comment thread unit_tests/test_cli_breakup.py Outdated
Comment thread unit_tests/test_cli_breakup.py Outdated


def _expected_masks(values, mapping):
"""Independent oracle: per-category membership masks via np.isin."""

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not sure how independent from

def land_cover_mask(ds_land_cover, mapped, category_id):
this really is, but fine for now.

Comment thread workflow/scripts/clip_and_rasterise_polys.py
irm-codebase and others added 2 commits September 24, 2026 13:09
* Improve unit test isolation.

Co-authored-by: Codex <codex@openai.com>

* Fix Windows bug on unit tests.

Root cause was attempt to delete open file.

---

Co-authored-by: Codex <codex@openai.com>

---------

Co-authored-by: Codex <codex@openai.com>
Co-authored-by: Ivan Ruiz Manuel <72193617+irm-codebase@users.noreply.github.com>
Co-authored-by: Jann Launer <32454596+jnnr@users.noreply.github.com>

@irm-codebase irm-codebase left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looking good!

@irm-codebase
irm-codebase merged commit 5f472e0 into main Sep 24, 2026
4 checks passed
@irm-codebase
irm-codebase deleted the feature-performance branch September 24, 2026 17:38
irm-codebase added a commit that referenced this pull request Sep 25, 2026
Feature performance

-----------
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Ivan Ruiz Manuel <72193617+irm-codebase@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

High memory consumption

3 participants