Skip to content

feat(mpmc): add capacity reservation to bounded senders - #334

Merged
BewareMyPower merged 3 commits into
apache:mainfrom
jiengup:add-capacity-reservation-to-bounded-mpmc-senders
Sep 29, 2026
Merged

BewareMyPower merged 3 commits into
apache:mainfrom
jiengup:add-capacity-reservation-to-bounded-mpmc-senders

Conversation

@jiengup

@jiengup jiengup commented Sep 26, 2026

Copy link
Copy Markdown
Contributor

Summary

Closes #297.

  • Add BoundedSender::reserve, try_reserve, and a borrowed mpmc::Permit so callers can wait for capacity before constructing a message, as requested in #297.
  • Add integration coverage for exact capacity, FIFO grants, cancellation, permit release, and receiver disconnection. Update the API documentation and changelog.
  • Commits since upstream/main: c42f549 (contract tests) and 28a0b14 (implementation).

Design Notes

Queued messages, held permits, and grants to waiting senders share the configured capacity. Sends and reservations receive grants in wait-queue order. Cancelling a granted operation or dropping an unused permit passes its slot to the next waiter. A permit reserves capacity, not message order, and Permit::send returns the value if receivers have disconnected.

Benchmarks

Lower is better. Results compare upstream/main (0f46831) with this branch (28a0b14). Each value is the median of 10 Divan run medians on an Apple M3 Pro with Rust 1.98.0. Microbenchmarks use capacity 1; batch benchmarks use capacity 64 and transfer 16,384 messages per iteration.

Existing bounded MPMC operation upstream/main This branch Time change
try_send → try_recv 11.47 ns 12.12 ns +5.7%
send → recv 15.56 ns 16.73 ns +7.5%
Wake blocked send 33.83 ns 38.39 ns +13.5%
Threads, producers→consumers 1→1 1,178.5 µs 1,164.5 µs −1.2%
Threads 1→8 16,050 µs 15,505 µs −3.4%
Threads 8→1 15,505 µs 14,355 µs −7.4%
Threads 8→8 5,959.5 µs 6,359 µs +6.7%
Tokio, current thread 1→1 268.7 µs 278.5 µs +3.6%
Tokio, current thread 1→8 336.2 µs 339.2 µs +0.9%
Tokio, current thread 8→1 301.3 µs 330.5 µs +9.7%
Tokio, current thread 8→8 307.2 µs 320.0 µs +4.2%
Tokio, 4 workers 1→1 385.0 µs 395.0 µs +2.6%
Tokio, 4 workers 1→8 1,099.5 µs 1,187.5 µs +8.0%
Tokio, 4 workers 8→1 1,095.5 µs 1,262.0 µs +15.2%
Tokio, 4 workers 8→8 1,318.0 µs 1,363.0 µs +3.4%
New reservation operation This branch
try_reserve → Permit::send → try_recv 16.52 ns
reserve → Permit::send → recv 21.56 ns
try_reserve → drop permit 10.17 ns
reserve → drop permit 12.35 ns
try_reserve on a full queue 4.64 ns
Wake blocked reserve, then send 31.40 ns

Existing single-thread round trips are 5.7%–13.5% slower. Batch results vary by topology and include both improvements and regressions; the largest measured regression is 15.2% for Tokio with 4 workers and 8 producers→1 consumer. Reservation-specific cases were measured in a temporary benchmark copy and are not part of this PR.

Validation

  • cargo x test passed, including the bounded MPMC reservation tests and doctests.

@jiengup
jiengup marked this pull request as ready for review September 26, 2026 15:39

@BewareMyPower BewareMyPower left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

From your benchmark, this PR introduces performance regressions in many cases, could you explain the reason?

@jiengup

jiengup commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor Author

@BewareMyPower

From your benchmark, this PR introduces performance regressions in many cases, could you explain the reason?

Larger Benchmarks

I reran the bounded MPMC benchmarks with larger samples, alternating upstream/main and this branch to reduce run-order effects. Each microbenchmark now uses 4,096,000 iterations per run (160× the previous count), across 20 paired runs. Times are lower-is-better; the percentage is the median of the paired ratios. The noticeable figure was bolded.

Existing bounded MPMC operation upstream/main median This branch median Paired time change Slower pairs
try_send → try_recv 12.03 ns 12.59 ns +4.7% 18/20
send → recv 16.94 ns 17.66 ns +5.7% 18/20
Wake a blocked send 35.07 ns 40.09 ns +14.8% 17/20
Threads, producers→consumers 1→1 1,173.0 µs 1,256.5 µs +2.1% 6/10
Threads 1→8 15,740.0 µs 16,180.0 µs +0.2% 5/10
Threads 8→1 15,625.0 µs 14,660.0 µs −0.6% 5/10
Threads 8→8 5,921.5 µs 6,336.0 µs +6.9% 9/10
Tokio, current thread 1→1 279.4 µs 285.7 µs +1.4% 6/10
Tokio, current thread 1→8 346.4 µs 349.1 µs +0.9% 7/10
Tokio, current thread 8→1 310.9 µs 334.7 µs +8.8% 9/10
Tokio, current thread 8→8 317.6 µs 329.1 µs +3.4% 7/10
Tokio, 4 workers 1→1 388.6 µs 404.8 µs +4.4% 10/10
Tokio, 4 workers 1→8 1,076.0 µs 1,051.5 µs +3.8% 6/10
Tokio, 4 workers 8→1 1,054.5 µs 1,045.0 µs +3.0% 6/10
Tokio, 4 workers 8→8 1,232.0 µs 1,162.0 µs −9.9% 2/10

I have updated that earlier claim from the PR description.

Conclusion

  • I believe most cases are due to delays introduced by additional code, or test jitter in a multi-threaded environment (e.g., Tokio, 4 workers 8→1’s +15.2% failed to replicate in a larger sample)
  • One of the significant costs comes from a big lock maintaining an additional variable available that is frequently accessed. Before this change, capacity could be checked against the queue length. Permits and grants to waiting senders now own slots without appearing in that queue, so a successful send decrements available and a receive releases or transfers that slot. In a temporary attribution build that skipped those writes, the matched try_send/try_recv round-trip gap fell from 14.3% to 2.4%; the send-only gap fell from 13.2% to approximately zero. This variable is key to express the reserve semantic.
  • The blocked-send result has a separate cause. Previously, the resumed send future could enqueue its value while holding the lock used to claim capacity. It now receives a permit from reserve() and calls Permit::send, which acquires the lock again to enqueue the value. The measured difference is about 5 ns per backpressure cycle; the roughly 15% figure is large because the entire benchmark takes only about 35–40 ns. This benchmark measures the queue transitions with a no-op waker, not executor scheduling. Currently, this implementation is aligned with mpsc, and they can be further optimized together, which is non-blocking for this PR.

@BewareMyPower

Copy link
Copy Markdown
Contributor

a big lock maintaining an additional variable available that is frequently accessed.

so is it necessary? or could you figure out a better approach?

@jiengup

jiengup commented Sep 29, 2026

Copy link
Copy Markdown
Contributor Author

a big lock maintaining an additional variable available that is frequently accessed.

so is it necessary? or could you figure out a better approach?

To follow the contract described in #334, I think it's necessary now, and the alignment with mpsc can help with the unified optimization in the future.

Do you think performance is in this PR's scope? Then I'll try to do it further.

@BewareMyPower

BewareMyPower commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

I think it's okay, we can improve the performance later if possible.

@BewareMyPower
BewareMyPower merged commit b383e93 into apache:main Sep 29, 2026
9 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add capacity reservation to bounded MPMC senders

2 participants