Skip to content

Fix flaky tests rel 2 stable - #2061

Draft
Alena0704 wants to merge 7 commits into
apache:REL_2_STABLEfrom
Alena0704:fix-flaky-tests-rel-2-stable
Draft

Alena0704 wants to merge 7 commits into
apache:REL_2_STABLEfrom
Alena0704:fix-flaky-tests-rel-2-stable

Conversation

@Alena0704

Copy link
Copy Markdown
Collaborator

Collect all fixes from #2032 that can improve CI stability after having reviewed with refactored commit messages.

Author: Dianjin Wang wangdianjin@gmail.com
Assisted-by: Claude Code

Type of Change

  • Bug fix (non-breaking change)
  • New feature (non-breaking change)
  • Breaking change (fix or feature with breaking changes)
  • Documentation update

Breaking Changes

Test Plan

  • Unit tests added/updated
  • Integration tests added/updated
  • Passed make installcheck
  • Passed make -C src/test installcheck-cbdb-parallel

Impact

Performance:

User-facing changes:

Dependencies:

Checklist

Additional Context

CI Skip Instructions


@Alena0704
Alena0704 force-pushed the fix-flaky-tests-rel-2-stable branch from 435867e to 335738d Compare September 29, 2026 16:58
@leborchuk leborchuk added the CI label Oct 2, 2026
@Alena0704
Alena0704 force-pushed the fix-flaky-tests-rel-2-stable branch from 335738d to 6a1a193 Compare October 8, 2026 09:17
@Alena0704 Alena0704 mentioned this pull request Oct 8, 2026
5 of 13 tasks
@Alena0704
Alena0704 requested review from leborchuk and tuhaihe October 8, 2026 09:19
@Alena0704

Alena0704 commented Oct 8, 2026 •

Copy link
Copy Markdown
Collaborator Author

This pull request is backport fixes of flacky tests for REL_2_STABLE from #2060

@Alena0704
Alena0704 force-pushed the fix-flaky-tests-rel-2-stable branch from d2f12e0 to 462a395 Compare October 8, 2026 09:56
@Alena0704
Alena0704 force-pushed the fix-flaky-tests-rel-2-stable branch from 462a395 to 0e06bc0 Compare October 8, 2026 18:07
@Alena0704

Alena0704 commented Oct 8, 2026 •

Copy link
Copy Markdown
Collaborator Author

CI failures revealed that the test fixes need adaptation for REL_2_STABLE. I'm reanalyzing commits and tests on REL_2_STABLE.

@Alena0704
Alena0704 marked this pull request as draft October 8, 2026 18:12
@Alena0704
Alena0704 force-pushed the fix-flaky-tests-rel-2-stable branch 3 times, most recently from 55722c0 to 7571857 Compare October 8, 2026 22:04
tuhaihe and others added 7 commits October 9, 2026 15:33
REL_2_STABLE has no SQL function to force pending table statistics to
flush. Close each writer session to flush on backend exit, then poll
the collector counters with a fresh statistics snapshot on every retry.
Restore on_change mode when reconnecting for the third scenario.

Wait for analyze, autoanalyze and modification counts before checking
the results, and raise an error if the expected state is not reached.
Apply the helper and expected results to regular, PAX and single-node
tests.

Assisted-by: Claude Code
Assisted-by: ChatGPT
Backpatch-through: REL_2_STABLE
Assisted-by: Claude Code
Backpatch-through: REL_2_STABLE
The test already averages CPU samples, but the expected values are
still too low. The suite sets the cluster CPU limit to 100%, so expect
100% for one uncapped group and 33%/67% for groups with weights 100/200.

Keep the existing tolerance of 10 percentage points.

Assisted-by: Claude Code
Backpatch-through: REL_2_STABLE
These tasks only check that valid schedules are accepted.
Deactivate them after creation to avoid background runs interfering
with cleanup and the oid_wraparound test.

Assisted-by: Claude Code
Backpatch-through: REL_2_STABLE
Close both aborted-insert sessions before polling statistics so pending
reports are flushed on backend exit. Wait on segment 1, which contains
all rows: pg_stat_all_tables aggregates its tuple counters from segments,
so waiting on the coordinator does not synchronize the following read.
Keep the statistics check before DELETE to avoid racing its counters.

Wait for the second VACUUM count instead of a dead-tuple count that was
already zero after the first VACUUM.

(cherry picked from commit fba73b0)

Assisted-by: ChatGPT
hot_standby/transaction_isolation deliberately panics the standby and
reconnects to check distributed transaction visibility after recovery.
The new connection can arrive before consistent recovery is reached,
when the server reports "the database system is not yet accepting
connections". This message was missing from the connection retry list.

Retry this transient state using the existing attempt limit and interval.
Keep the crash scenario and all transaction visibility checks intact.

Backport required: REL_2_STABLE is affected by the same race.

(cherry picked from commit 9ec9f4c)
The final termination scenario can receive a disconnect before backend
exit has removed its temporary table. Keep pg_terminate_backend()
asynchronous and use wait_for_temp_table_cleanup() to poll pg_class
until no temporary relations belong to the test role, before DROP ROLE.
This prevents a leftover role and resource group from affecting later
tests. Bound the retries and report an error if cleanup never completes.
@Alena0704
Alena0704 force-pushed the fix-flaky-tests-rel-2-stable branch from 7571857 to 8f80d90 Compare October 9, 2026 12:33
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants