Skip to content

Parallelize symmetric sorting comparisons - #4810

Open
JESUSROYETH wants to merge 1 commit into
SpikeInterface:mainfrom
JESUSROYETH:radar/perf-parallel-symmetric-comparison
Open

JESUSROYETH wants to merge 1 commit into
SpikeInterface:mainfrom
JESUSROYETH:radar/perf-parallel-symmetric-comparison

Conversation

@JESUSROYETH

Copy link
Copy Markdown
Contributor

What this changes

Symmetric sorting comparisons with the default count agreement method run the forward and reverse spike-matching scans one after the other. The two scans are independent, and the existing Numba kernel already releases the GIL, so nothing stops them running at the same time. (The distance agreement method already computes both directions in one pass, so it's unaffected.)

This adds n_jobs to compare_two_sorters() and the underlying agreement helpers, at the end of each signature. With n_jobs>1, the reverse scan runs in a worker thread while the caller runs the forward scan. The default stays at 1: for very small sortings the thread setup costs more than it saves, and with only two directions, two jobs is the useful ceiling anyway. n_jobs still goes through fix_job_kwargs(), but it doesn't route through TimeSeriesChunkExecutor — there's no time-series chunking here, just two independent whole-input scans, so it only needs one executor worker plus the caller. GroundTruthComparison uses ensure_symmetry=False, so it's unaffected. MultiSortingComparison doesn't pass n_jobs through either: it already parallelises across sorter pairs, and nesting a thread pool inside each worker process would oversubscribe the machine. So today's gain is scoped to a direct compare_two_sorters()/SymmetricSortingComparison call.

Benchmark

I tested the complete compare_two_sorters() call on a published Kilosort4 sorting with 713 units / 11.69M spikes against its 650-unit curated view / 10.60M spikes. Results are 7 alternating runs on an Intel i9-13900HX, Python 3.12.3, NumPy 2.3.5 and Numba 0.67.0:

median [range] change
n_jobs=1 1.189 s [1.154, 1.194] baseline
n_jobs=2 0.770 s [0.755, 0.789] 35.3% lower latency, 0.42 s saved

Each measured run produced the same SHA256 over match counts, agreement scores and both Hungarian mappings.

Validation

  • New serial/threaded equivalence tests, including a multi-segment sorting compared on match counts, agreement scores and both Hungarian mappings. Confirmed to fail on unpatched main and pass with this change.
  • Comparison tests: 25 passed. The suite's pre-existing compare_multiple_sorters process-parallelism timing test, unrelated to this change, also passed separately.
  • In-memory, NumpyFolder and zarr-backed sortings checked.
  • Black, whitespace, patch-apply and SpikeInterface review-pattern checks are clean.
  • Benchmarked and tested on Linux only. ThreadPoolExecutor and the underlying nogil=True Numba kernel are available on macOS/Windows too, but I haven't verified this path there.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant