Skip to content

Write the per-tile final_cat as HDF5, one dataset per column - #942

Open
cailmdaley wants to merge 3 commits into
developfrom
feat/final-cat-hdf5
Open

cailmdaley wants to merge 3 commits into
developfrom
feat/final-cat-hdf5

Conversation

@cailmdaley

@cailmdaley cailmdaley commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Closes #941. Builds on #939 and #937 (both merged).

Change

  • make_cat writes the per-tile catalogue as HDF5, in one pass. The runner collects every stage's columns (SExtractor, TILE_ID/TILE_UNIQUE_ID, ngmix, per-epoch PSF slots, MASK_*) in one ordered dict and writes final_cat<num>.hdf5. Each column is one lzf dataset; vector columns are 2-D. Columns keep stage order and dtype (byte order normalised to native). The file has no attributes, because the FITS file carried nothing beyond its column cards and EXTNAME.
  • WORK_DIR is removed (runner and config_tile_Mc.ini). With a single write there is nothing to stage: a real 9,226 × 174 tile writes in 0.17 s to NFS on candide.
  • Readers. create_final_cat.read_data reads per-column HDF5 into the structured array the merge already writes, and still reads FITS. find_final_cat prefers .hdf5 over .fits in both layouts. FITS support stays for runs already on disk and for sp_validation's image-sims im_merge rule. merge_final_cat reads final_cat-<ID>.hdf5 and drops its unused --hdu.
  • Workflow. final_cat(), tile_make_cat and astra.yaml declare final_cat-<tile>.hdf5.

The merged final_cat_<run>.hdf5 layout does not change (one compound dataset per tile).

Campaign boundary

The declared per-tile output and the tile_make_cat params pin both change. A resumed campaign finds no final_cat-<tile>.hdf5 and re-declares every tile, so start a fresh root.

Schema note

The FITS writer promoted every float to float64. HDF5 keeps SExtractor's float32, so in SExtractor-mode (image-sims) campaigns 12 merged columns become float32 with identical values (MAG_*, FLUX_*, FLUXERR_*, MAGERR_*, FLUX_RADIUS, SNR_WIN, FWHM_*). A merged row shrinks from 530 to 482 bytes. Data campaigns keep their dtypes: DR6 detection columns are already f8/i4/i8 (checked on a DR6 sexcat).

Verification

  • Real tile (image-sims 1z2z_grid_3, 273.282): make_cat as FITS (make_cat: write the final catalogue once per save stage, node-local #939) and as HDF5 (this branch) gives the same 174 columns in the same order with equal values (NaN-aware); only the float32 dtypes differ. Merging each through read_data + copy_data gives identical rows.
  • Tests (dev SIF shapepipe_develop-dev-20260926, candide): pytest tests -m "not slow" gives 861 passed, 1 skipped, including new tests for the per-column lzf writer, merge invariants on make_cat's writer, and read_data HDF5/FITS parity.
  • uvx astra-tools@0.2.17 validate passes.

Not changed

The Gen-2 helpers scripts/python/merge_final_cat.py, get_number_objects.py, combine_runs.bash, post_proc_sp.bash and example/unions_800/config_tile_match_ext_r.ini still expect FITS. sp_validation reads only the merged file and needs no change.

Claude Opus 5.5 on behalf of Cail

🤖 Generated with Claude Code

https://claude.ai/code/session_01JkM3cmXPuuSrNAf7vJde9e

cailmdaley and others added 3 commits October 2, 2026 12:25
make_cat collects every stage's columns (SExtractor, TILE_ID and
TILE_UNIQUE_ID, ngmix, per-epoch PSF slots, MASK_*) in one ordered dict
and writes final_cat<num>.hdf5 at the end: one lzf-compressed dataset per
column, vector columns as 2-D datasets, native dtypes and byte order.

With a single write, building on node-local disk has nothing left to
buy (the write of a 9,226 x 174 tile takes 0.17 s on candide's NFS), so
WORK_DIR and its stale-file handling go.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JkM3cmXPuuSrNAf7vJde9e
create_final_cat.read_data reads a per-tile HDF5 catalogue (one dataset
per column) into the same structured array the merge already writes,
and still reads FITS catalogues made before the format change, so hand
merges over existing runs keep working. find_final_cat looks for .hdf5
before .fits in both layouts. merge_final_cat reads
tiles/<prefix>/<ID>/final_cat-<ID>.hdf5 and drops the unused --hdu.

The merge invariants now write their tile catalogues with make_cat's own
writer, cover a vector column, and check that read_data gives one result
from the HDF5 and FITS forms of a catalogue.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JkM3cmXPuuSrNAf7vJde9e
tile_make_cat publishes final_cat-<tile>.hdf5 and final_cat_merge reads
it. The rendered tile_make_cat shell changes, so the params pin moves:
this is a campaign boundary. astra.yaml records the format.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JkM3cmXPuuSrNAf7vJde9e

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Write the per-tile final_cat as HDF5, matching the campaign merge

1 participant