Skip to content

fix(sqlserver): record transitions before the claim and the report take X locks - #11

Merged
pdevito3 merged 7 commits into
mainfrom
fix/sqlserver-relinquish-deadlock-flake
Oct 7, 2026
Merged

pdevito3 merged 7 commits into
mainfrom
fix/sqlserver-relinquish-deadlock-flake

Conversation

@pdevito3

@pdevito3 pdevito3 commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Two SQL Server tests can fail with absorbed deadlocks (error 1205). Both tests pin that count to 0.

  • Concurrent_relinquishes_and_claims_never_deadlock failed on CI on main.
  • Concurrent_outcome_reports_never_deadlock failed on CI for the first commit of this PR. On main, it fails in 6 runs out of 6 when it runs alone.

The claim and the report took X locks on their jobs first, then wrote the Transition Log. The FK check of the transition INSERT can scan jobs, and the deadlock graphs show that scan wait for S locks on the rows of a concurrent claim or report. Each side holds X locks that the other side scans, so SQL Server kills one side.

The claim and the report now write the Transition Log before the UPDATE, in the same batch:

 claim / report batch
-  UPDATE jobs (X locks)  OUTPUT the rows
-  INSERT transitions     -- FK scan waits on rival X locks
+  INSERT @batch  FROM jobs WITH (UPDLOCK)   -- U locks only; the report applies the Effect-Once fence here
+  INSERT transitions FROM @batch            -- U is compatible with the S of the FK scan
+  UPDATE jobs FROM @batch                   -- X locks come last

The expiry and relinquish paths already obey this order. All batch writers (claim, report, expiry, relinquish) now share one transition INSERT (InsertTransitionsFromBatch) and one prune step (PruneTransitionsAsync). The single-row RecordTransitionAsync paths (enqueue, cancel, requeue, cascades, schedule mint) still do the UPDATE first.

Each path also uses one less round trip:

Budget Before After
Claim of 32 jobs 6 5
Report of 32 rows 3 2
Report of 32 rows with job output 35 34

Test changes

  • Concurrent_outcome_reports_never_deadlock now runs ALTER DATABASE SCOPED CONFIGURATION CLEAR PROCEDURE_CACHE first. Before, an earlier test in the run could leave a narrow plan in the cache, and the test then passed on any lock order.
  • The poison is now one 96-row report, not an expiry sweep. Each writer has its own batch text and its own cached plan, so only a wide report poisons the report plan. The test claims the 96 jobs in passes, because one claim returns at most Bounds.MaxClaimBatch (32) rows. The old poison claimed only 32.
  • A_claim_writes_its_transitions_before_it_takes_X_locks and A_report_writes_its_transitions_before_it_takes_X_locks pin the lock order with no race. A rival session holds an S lock on one row of the batch. The store passes it with U, inserts the transitions, then waits at the UPDATE. A dirty read taken during the wait must show all 4 entries.

Evidence

Local SQL Server 2022 container, limited to 2 CPUs.

  • Before (main, and the first commit of this PR):

    • Relinquish test: failed.
    • Report test, alone: failed 6 runs out of 6 on main, and 7 out of 8 after the first commit.
  • After (f00a23c):

    Run Result
    Concurrent_outcome_reports_never_deadlock, alone 8 out of 8 passed
    Concurrent_relinquishes_and_claims_never_deadlock, alone (both history policies) 10 out of 10 passed
    SqlServerConcurrentMaintenanceTests (10 tests) 5 out of 5 passed
    All 263 tests in BackWave.SqlServer.Tests 2 out of 2 passed

    The build has 0 warnings.

  • After the merge of main (with feat(monitor): add a filtered job count #9 and feat: record why a job went back to Scheduled and show Retrying jobs #10): all 273 tests in BackWave.SqlServer.Tests passed, and SqlServerConcurrentMaintenanceTests passed 3 runs out of 3. The claim OUTPUT now also returns retry_cause.

Merge Danger

Door: two-way

The change is SQL text inside SqlServerJobStore. It has no schema change and no change to the stored data, so a revert puts the old statements back.

Blast Radius: SQL Server claim and report

Only the SQL Server adapter changes. Every claim and every outcome report on SQL Server runs the new batch, so a fault here stops work on SQL Server fleets. The Transition Log entries keep the same content.

Concurrent_relinquishes_and_claims_never_deadlock failed on CI with an
absorbed deadlock. The deadlock graphs show claim against claim. Each
claim held X locks on its leased jobs and waited for S locks inside the
FK check of its Transition Log INSERT. When the plan scans jobs, that
check reads the rows that other claimers hold.

The claim now writes its Transition Log entries in the same batch as
the lease, before the UPDATE. At that point the candidates hold only U
locks, and U locks are compatible with S locks. The expiry and
relinquish paths already obey this rule.

The claim also uses one less round trip. The budget for a claim of 32
jobs decreases from 6 statements to 5.
… locks

Concurrent_outcome_reports_never_deadlock failed on CI for this branch
with 19 absorbed deadlocks. The fault is older than the claim fix. On
main, the test fails in 6 runs out of 6 when it runs alone, with 9
absorbed deadlocks each time. In the full run, the tests before it
change the plan cache, so CI usually passes.

The report wrote its outcomes first and its Transition Log entries
second. The FK check of the transition INSERT can scan jobs, and its S
locks wait behind the X locks of concurrent reporters.

The report now uses the same order as the claim. One batch locks the
fenced rows with UPDLOCK into a table variable, writes the transitions
while the rows hold only U locks, and then updates the rows. The U
locks keep the fence valid until the UPDATE.

The report also uses one less round trip. The drain budget decreases
from 3 statements to 2, and the budget with job output from 35 to 34.
@pdevito3 pdevito3 changed the title fix(sqlserver): record claim transitions before the lease takes X locks fix(sqlserver): record transitions before the claim and the report take X locks Oct 6, 2026
SQL Server caches one plan per batch text. The report now writes its
transitions in its own batch, so the 96-row lease sweep no longer
changed the plan of the report, and the test passed on any lock order.

Clear the plan cache, then compile the report plan with a 96-row
report. With X locks taken before the transition INSERT, the test now
fails 3 runs out of 3. The old setup passed 3 out of 3 on that order.
The deadlock tests for the claim and the report are races, so they
catch the old lock order only when the plan loses the race.

Two new tests hold an S lock on one batch row in a rival session. U
passes the S lock and X does not, so the store blocks at its UPDATE. A
dirty read then counts the transition entries that the blocked store
already wrote. With the store from main, both tests see 0 of 4 entries
and fail on every run.
The report no longer goes through the batch recorder, the prune payload
is not always TransitionRow JSON, and one method named its cap twice.
…der note

Replace the two repeated deadlock explanations with the one-line
reference that the relinquish uses, and document the @Batch columns
that InsertTransitionsFromBatch reads.
…ish-deadlock-flake

# Conflicts:
#	src/BackWave.SqlServer/SqlServerJobStore.cs
@pdevito3
pdevito3 merged commit 93af59f into main Oct 7, 2026
2 checks passed
@pdevito3
pdevito3 deleted the fix/sqlserver-relinquish-deadlock-flake branch October 7, 2026 21:00
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant