Add compressed binary accessibility caches (.agz) - #250
Conversation
|
@martin-raden The implementation for #245 is ready for review. This PR adds compressed |
|
|
||
| // gzipped input file stream | ||
| if (boost::iends_with(in, ".gz")) { | ||
| if (boost::iends_with(in, ".gz") || boost::iends_with(in, ".agz")) { |
There was a problem hiding this comment.
check if useful for this kind of data
|
revise the code of this PR given the following requests:
|
|
@martin-raden Addressed both requested steps in 2e2c5f0, on top of your commits:
Full measurements, direct/generic comparisons, raw trial data and reproduction commands are included. All 48 benchmark reloads matched exactly. Release/native and debug/Kokkos suites pass (73,630 assertions in 46 API cases, plus both CLI suites), along with installed-header and original-archive compatibility checks. |
Repeated target screens spend time formatting and parsing large accessibility matrices. This adds gzip-compressed
.agzcaches containing versioned Boost binary archives of exact internal ED values, including the extra interval length needed for dangling-end probabilities.Use
--out=tAcc:target.agz(or any existing query/target Acc/Pu output), then reuse it with--tAcc=E --tAccFile=target.agz. The extension selects the binary reader in either E or P mode; both use the stored ED values unchanged. Existing text and.gzbehavior is preserved.Unconstrained
AccessibilityVrnaandAccessibilityFromStreampass their stored ED rows directly to Boost through non-owning row views, without copying rows or callinggetED()per cell. Other accessibility implementations and constrained data retain generic export to preserve their accessibility semantics. Loading fills retained matrix rows in place and validates discarded tails when reading a narrower band. The version 1 archive layout is unchanged.The reader checks the format, sequence, dimensions, energy values, trailing data and gzip integrity. Smaller requested interaction lengths are supported. Boost.Serialization is checked during configure and linked into the CLI and pkg-config consumers. README, CLI help and ChangeLog document usage and compatibility.
Validation on Linux x86-64 with GCC 16.2.0, Boost 1.85.0 and ViennaRNA 2.7.2:
std::mdspanand debug with bundled Kokkosmdspan:make tests -j2passes all 73,630 assertions in 46 API cases, plus both CLI suites. The debug build was built from the source distribution.getED()calls and byte-identical generic/direct payloads; a constrained short-sequence regression preserves masking. Invalid discarded-band values are rejected.make install, independent compilation of affected installed public headers, and a linked installed-library consumer/benchmark. An archive produced by the original version 1 implementation reloads and re-exports byte-for-byte identically.make distincludes the benchmark source, documentation and per-trial CSV data;git diff --checkpasses.Compression was measured after the direct matrix optimization, using the same 100,000-base ED matrix for each format. Median of three trials for
AccessibilityVrna:Gzip is retained: it costs roughly 2.7–2.9 seconds per write and 0.23–0.25 seconds per read in this run, but saves 58–59% of storage. Raw archives are about 2.4 times larger. The benchmark also compares direct/generic export and
AccessibilityFromStream; all 48 reloads matched every ED cell exactly. Gzip dominates compressed write time, while the raw timings show the benefit of direct export.Reproduction commands, hardware, full tables and CSV data are included. Timings include stream close but exclude folding/verification, use cached filesystem reads, and do not include
fsync. They measure local I/O, not end-to-end genome-screen speedups.Compatibility: archives use Boost's native binary format, requiring a compatible architecture/archive version. Cached temperature/folding/model settings are not stored or checked; reuse with matching settings. macOS/Apple Clang has not been tested locally.
Fixes #245.