Write the per-tile final_cat as HDF5, one dataset per column - #942
Open
cailmdaley wants to merge 3 commits into
Open
cailmdaley wants to merge 3 commits into
cailmdaley wants to merge 3 commits into
Conversation
make_cat collects every stage's columns (SExtractor, TILE_ID and TILE_UNIQUE_ID, ngmix, per-epoch PSF slots, MASK_*) in one ordered dict and writes final_cat<num>.hdf5 at the end: one lzf-compressed dataset per column, vector columns as 2-D datasets, native dtypes and byte order. With a single write, building on node-local disk has nothing left to buy (the write of a 9,226 x 174 tile takes 0.17 s on candide's NFS), so WORK_DIR and its stale-file handling go. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JkM3cmXPuuSrNAf7vJde9e
create_final_cat.read_data reads a per-tile HDF5 catalogue (one dataset per column) into the same structured array the merge already writes, and still reads FITS catalogues made before the format change, so hand merges over existing runs keep working. find_final_cat looks for .hdf5 before .fits in both layouts. merge_final_cat reads tiles/<prefix>/<ID>/final_cat-<ID>.hdf5 and drops the unused --hdu. The merge invariants now write their tile catalogues with make_cat's own writer, cover a vector column, and check that read_data gives one result from the HDF5 and FITS forms of a catalogue. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JkM3cmXPuuSrNAf7vJde9e
tile_make_cat publishes final_cat-<tile>.hdf5 and final_cat_merge reads it. The rendered tile_make_cat shell changes, so the params pin moves: this is a campaign boundary. astra.yaml records the format. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JkM3cmXPuuSrNAf7vJde9e
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #941. Builds on #939 and #937 (both merged).
Change
TILE_ID/TILE_UNIQUE_ID, ngmix, per-epoch PSF slots,MASK_*) in one ordered dict and writesfinal_cat<num>.hdf5. Each column is one lzf dataset; vector columns are 2-D. Columns keep stage order and dtype (byte order normalised to native). The file has no attributes, because the FITS file carried nothing beyond its column cards andEXTNAME.WORK_DIRis removed (runner andconfig_tile_Mc.ini). With a single write there is nothing to stage: a real 9,226 × 174 tile writes in 0.17 s to NFS on candide.create_final_cat.read_datareads per-column HDF5 into the structured array the merge already writes, and still reads FITS.find_final_catprefers.hdf5over.fitsin both layouts. FITS support stays for runs already on disk and for sp_validation's image-simsim_mergerule.merge_final_catreadsfinal_cat-<ID>.hdf5and drops its unused--hdu.final_cat(),tile_make_catand astra.yaml declarefinal_cat-<tile>.hdf5.The merged
final_cat_<run>.hdf5layout does not change (one compound dataset per tile).Campaign boundary
The declared per-tile output and the
tile_make_catparams pin both change. A resumed campaign finds nofinal_cat-<tile>.hdf5and re-declares every tile, so start a fresh root.Schema note
The FITS writer promoted every float to float64. HDF5 keeps SExtractor's float32, so in SExtractor-mode (image-sims) campaigns 12 merged columns become float32 with identical values (
MAG_*,FLUX_*,FLUXERR_*,MAGERR_*,FLUX_RADIUS,SNR_WIN,FWHM_*). A merged row shrinks from 530 to 482 bytes. Data campaigns keep their dtypes: DR6 detection columns are already f8/i4/i8 (checked on a DR6 sexcat).Verification
read_data+copy_datagives identical rows.shapepipe_develop-dev-20260926, candide):pytest tests -m "not slow"gives 861 passed, 1 skipped, including new tests for the per-column lzf writer, merge invariants on make_cat's writer, andread_dataHDF5/FITS parity.uvx astra-tools@0.2.17 validatepasses.Not changed
The Gen-2 helpers
scripts/python/merge_final_cat.py,get_number_objects.py,combine_runs.bash,post_proc_sp.bashandexample/unions_800/config_tile_match_ext_r.inistill expect FITS. sp_validation reads only the merged file and needs no change.Claude Opus 5.5 on behalf of Cail
🤖 Generated with Claude Code
https://claude.ai/code/session_01JkM3cmXPuuSrNAf7vJde9e