Repository navigation
Conversation
01237c9 to
3b56d29
Compare
c2a5026 to
d8fed9e
Compare
c999d4e to
d3a10f3
Compare
14db4ac to
bad9884
Compare
A HOT-indexed update plants index entries that point at mid-chain heap-only tuples, so a dead chain member cannot simply be removed: a not-yet-swept index entry may still arrive at it, and the per-hop modified-attrs bitmap on it is what a reader unions to judge staleness. Teach prune to collapse a dead chain prefix into xid-free forwarding stubs: each preserved dead key tuple is rewritten in place to a stub (frozen, natts == 0, HEAP_INDEXED_UPDATED, forwarding via t_ctid.offnum) that keeps its segment's modified-attrs bitmap, and a member whose attributes are wholly subsumed by later hops is reclaimed instead. Readers step through stubs transparently and still cross every surviving hop's bitmap. The collapse back to classic HOT is driven by prune: once a chain is fully dead, a later prune (heap_prune_chain / heap_prune_chain_find_live) reclaims its members and re-points the root redirect straight at the first live tuple. VACUUM's index cleanup sweeps the stale leaves; its second pass (lazy_vacuum_heap_page) does the usual LP_DEAD -> LP_UNUSED conversion and leaves the HOT-indexed collapse to prune. The collapse reuses the existing prune/freeze WAL via an xlhp_prune_items sub-record carrying the (offset, forward) stub pairs; no new record type is introduced. A page that still carries a preserved stub (or a redirect that forwards into a live HOT-indexed member) is kept non-all-visible so index-only scans heap-fetch through the chain; heap_page_would_be_all_visible recognizes both the redirect-to-SIU and the stub case explicitly. When prune does set a page all-visible, those carve-outs guarantee that no live HOT-indexed row on it still has index entries that disagree on offset, so it also clears VISIBILITYMAP_LOCATOR_SPLIT, in the same critical section and on the same prune/freeze record; redo clears it too. Co-authored-by: Greg Burd <greg@burd.me> Co-authored-by: Nathan Bossart <nathandbossart@gmail.com>
Expose the HOT-indexed activity counters maintained by the write path: pg_stat_all_tables.n_tup_hot_indexed_upd, the per-index n_tup_hot_indexed_upd_matched / n_tup_hot_indexed_upd_skipped counters in pg_stat_all_indexes, and pg_relation_hot_indexed_stats() reporting per-relation HOT-indexed chain composition. Document them in monitoring.sgml and the README. With statistics, prune/collapse, and amcheck recognition all in place, add the full feature test suite, which uses those facilities to verify behavior: - hot_indexed_updates (regression): eligibility and classification; selective maintenance across multiple/composite indexes; the crossed-attribute read path for equality, range, and inequality scans; a key cycled away and back (ABA), including across two distinct live rows; TOASTed indexed columns; partial-index predicate flips (key and non-key predicate columns); trigger-modified indexed columns; exclusion-constraint tables; partitioned tables; non-btree access methods (hash, GIN, GiST); a UNIQUE index on a type where image equality differs from operator equality; CREATE INDEX / REINDEX and DROP INDEX over live chains; prune reclamation, stub mixes, and re-collapse across partial VACUUMs; the never-all-visible guard; and DDL after a chain exists (ADD COLUMN crossing a bitmap-size boundary, DROP COLUMN). - hot_indexed_adversarial (isolation): concurrent UPDATE / VACUUM / prune and index scans, key cycling, aborts, and reader consistency across a concurrent collapse. - 054_hot_indexed_recovery (recovery): WAL replay of the chain and its collapse under wal_consistency_checking. - pg_surgery handling of HOT-indexed tuples and collapse-survivor stubs. Authored-by: Greg Burd <greg@burd.me>
verify_heapam must not flag the HOT-indexed artifacts as corruption: a live HEAP_INDEXED_UPDATED heap-only tuple whose mid-chain line pointer is preserved because an index entry still points at it, an xid-free collapse-survivor stub, and more than one LP_REDIRECT forwarding to the same live tuple are all legitimate. Recognize them and continue checking the rest of the chain. Cover this with an amcheck regression test, and add a pg_upgrade test that carries a relation with HOT-indexed chains, an ABA-cycled indexed column, an out-of-line indexed column, and VACUUM-collapsed stubs across an upgrade, verifying the data, verify_heapam, bt_index_check, and the chain scans on the new cluster. Authored-by: Greg Burd <greg@burd.me>
…dates Core disqualifies HOT-indexed updates unconditionally on the logical-replication apply path: HeapUpdateHotAllowable returns HEAP_UPDATE_ALL_INDEXES whenever IsLogicalWorker() is true, because a HOT-indexed update of a replica-identity attribute leaves a stale index leaf that the apply worker's replica-identity lookups would otherwise be exposed to. This commit relaxes that unconditional bail-out with a per-subscription hot_indexed_on_apply option (subhotindexedonapply: off / subset_only (default) / always). HeapUpdateHotAllowable now consults it when running in an apply worker, comparing the relation's indexed-attribute set against its primary-key attributes: "off" disqualifies HOT-indexed whenever any indexed attribute lies outside the primary key, "subset_only" requires the indexed attributes to be a subset of the primary key, and "always" applies no apply-path gating. The apply worker's replica-identity lookups (see RelationFindReplTupleByIndex) cope with the stale leaf, but only when the indexed attributes are a subset of the replica identity, which is what the default subset_only enforces. Wire the option through CREATE/ALTER SUBSCRIPTION, pg_subscription, pg_dump, and psql's \dRs+, and document it (create_subscription, alter_subscription, catalogs). Cover apply under each mode (039), apply under REPLICA IDENTITY FULL and a non-PK USING INDEX whose key is cycled (040), and decoding of HOT-indexed update chains (test_decoding). Authored-by: Greg Burd <greg@burd.me>
A/B and single-variant benchmark scripts for HOT-indexed updates: build two postgres variants, run pgbench workloads exercising classic-HOT, non-HOT, and HOT-indexed paths, and a self-contained bloat probe that reports the skip count (index writes avoided on unchanged indexes) and changed-index bounding. Not for merge; kept for evaluating the feature.
There was a problem hiding this comment.
🔍 OCR found 90 issue(s).
- 25 inline, 65 in summary (inline capped at 25)
📄 src/include/catalog/catversion.h
Remove this CATALOG_VERSION_NO bump. Bumping catversion is the committer's job at push time; including it in the patch is a top author mistake in catalog patches and causes needless merge conflicts against every other in-flight catalog change. (high confidence)
📄 src/backend/access/heap/hot_indexed_stats.c (L73-L80)
This function interprets every page as a heap page (PageGetItemId, HeapTupleHeader, HEAP_INDEXED_UPDATED, HotIndexedHeaderIsStub) but accepts any RELKIND_RELATION/MATVIEW/TOASTVALUE regardless of table AM. For a relation using a non-heap table AM, reading MAIN_FORKNUM pages and decoding them as heap tuples is undefined behavior (garbage results or crash). Add a heap-AM guard, as repack.c and heapam.c do, e.g. if (rel->rd_tableam != GetHeapamTableAmRoutine()) ereport(ERROR, ...).
📄 src/backend/access/heap/heapam_indexscan.c (L19-L22)
Duplicate #include "access/hot_indexed.h" (also included on the line above access/relscan.h). Remove the duplicate and keep a single include in correct alphabetical position (before access/relscan.h). The include guard hides it at compile time, but it is a defect and breaks the project's alphabetical include ordering.
💡 Suggested change
Before:
#include "access/hot_indexed.h"
#include "access/relscan.h"
#include "access/hot_indexed.h"
#include "access/sysattr.h"
After:
#include "access/hot_indexed.h"
#include "access/relscan.h"
#include "access/sysattr.h"
📄 src/backend/catalog/indexing.c (L146-L147)
The previous HOT-only skip path carried Assert(!ReindexIsProcessingIndex(RelationGetRelid(index))), a can't-happen invariant verifying that a classic-HOT catalog update never skips an index currently being reindexed. The runtime skip behavior here is equivalent and correct, but that defensive assertion is now gone entirely. Consider restoring it in the skip branch to preserve the invariant check. (moderate confidence, low severity)
💡 Suggested change
Before:
if (index_unchanged && !indexInfo->ii_Summarizing)
continue;
After:
if (index_unchanged && !indexInfo->ii_Summarizing)
{
Assert(!ReindexIsProcessingIndex(RelationGetRelid(index)));
continue;
}
📄 src/backend/executor/execTuples.c (L69-L69)
Orphan include: utils/datum.h declares datumGetSize/datumCopy/datumIsEqual/datum_image_eq etc., none of which are used in this file (the DatumGet*/PointerGetDatum macros come from postgres.h). This is unrelated diff churn that violates the minimal-diff discipline; drop it unless a symbol from datum.h is actually referenced.
📄 src/backend/executor/nodeModifyTable.c (L2318-L2320)
Per-row Bitmapset leak on the concurrent-update retry path. On the first ExecUpdateAct call, line 2444 sets updateCxt->modified_attrs to a palloc'd bitmap (query context). If table_tuple_update returns TM_Updated, ExecUpdate takes goto redo_act and re-enters ExecUpdateAct, which here resets modified_attrs = NULL without freeing the previous allocation. ExecUpdateEpilogue (the only bms_free site) never ran, because it's reached only after TM_Ok. The same leak occurs in the MERGE path (lmerge_matched), where ExecUpdateEpilogue is gated on result == TM_Ok and a concurrent update loops back through ExecUpdateAct. heapam_modified_attrs returns a non-NULL (possibly empty) allocated set, so this leaks one bitmap per retry for the statement's lifetime. Free before overwriting.
💡 Suggested change
Before:
/* Reset any state left over from a previous call */
updateCxt->modified_attrs = NULL;
updateCxt->row_moved = false;
After:
/* Reset any state left over from a previous call */
bms_free(updateCxt->modified_attrs);
updateCxt->modified_attrs = NULL;
updateCxt->row_moved = false;
📄 src/backend/replication/logical/worker.c (L6186-L6187)
The header comment says the mode is "cached", but the body comment (and the code) explicitly state the value is derived directly from MySubscription "rather than caching". These two comments contradict each other. Since the function reads MySubscription->hotindexedonapply on every call, drop the "cached" wording so the comment describes current behavior.
Confidence: high.
💡 Suggested change
Before:
* Return the cached HOT-indexed apply mode of the current logical replication
* worker's subscription.
After:
* Return the HOT-indexed apply mode of the current logical replication
* worker's subscription.
📄 src/include/replication/logicalworker.h (L28-L29)
Comment says "cached hot_indexed_on_apply mode", but the implementation in worker.c derives the value directly from MySubscription on each call (its own body comment notes it does so "rather than caching"). Remove "cached" to keep the comment consistent with the implementation.
Confidence: high.
💡 Suggested change
Before:
* Accessor for the cached hot_indexed_on_apply mode of the current apply
* worker's subscription. Returns a LOGICALREP_HOT_INDEXED_* code (see
After:
* Accessor for the hot_indexed_on_apply mode of the current apply
* worker's subscription. Returns a LOGICALREP_HOT_INDEXED_* code (see
📄 src/include/pgstat.h (L288-L288)
The on-disk/shared-memory stats structs PgStat_StatTabEntry and PgStat_StatIdxEntry gain new counters in this diff, changing their layout. pgstat_write_statsfile()/read path serialize these structs raw via shared_data_len, so PGSTAT_FILE_FORMAT_ID must be bumped whenever these structures change (per the comment just above: "PGSTAT_FILE_FORMAT_ID should be changed whenever any of these data structures change"). Without the bump, a stats file written by a pre-patch binary would be read with the wrong byte offsets. Unlike catversion.h, this ID is author-owned and expected to be bumped in the patch. (high confidence)
📄 src/include/pgstat.h (L858-L867)
These two counters live in the idx arm of the union shared with tab (PgStat_StatCounts). The neighboring index-only macros (pgstat_count_index_scan, pgstat_count_index_tuples) guard the idx.* write with Assert((rel)->pgstat_info->kind == PGSTAT_KIND_INDEX) precisely to catch accidental use on a non-index relation (which would corrupt the overlapping tab member). For consistency and safety, add the same assert here. (moderate confidence)
💡 Suggested change
Before:
#define pgstat_count_hot_indexed_upd_skipped(rel) \
do { \
if (pgstat_should_count_relation(rel)) \
(rel)->pgstat_info->idx.tuples_hot_indexed_upd_skipped++;\
} while (0)
#define pgstat_count_hot_indexed_upd_matched(rel) \
do { \
if (pgstat_should_count_relation(rel)) \
(rel)->pgstat_info->idx.tuples_hot_indexed_upd_matched++;\
} while (0)
After:
#define pgstat_count_hot_indexed_upd_skipped(rel) \
do { \
if (pgstat_should_count_relation(rel)) \
{ \
Assert((rel)->pgstat_info->kind == PGSTAT_KIND_INDEX); \
(rel)->pgstat_info->idx.tuples_hot_indexed_upd_skipped++; \
} \
} while (0)
📄 src/include/pgstat.h (L863-L863)
Backslash continuation alignment is off: the macro-name line here uses an extra tab compared to surrounding macros, and the body-line backslashes butt directly against the code with no alignment. This will not survive pgindent cleanly. (low confidence)
📄 src/include/catalog/pg_subscription.h (L180-L180)
pgindent aligns trailing field comments; this line uses a single space before /* while adjacent members (conflictlogrelid, slotname) use a tab to reach the comment column. Running pgindent will reformat this, so align it now to keep git diff --check / pgindent clean.
Confidence: low (cosmetic; pgindent would auto-fix).
📄 src/backend/executor/execIndexing.c (L1164-L1167)
When row_moved is true, ii_IndexNeedsUpdate is unconditionally true (|| short-circuits), so the result of bms_overlap is unused — yet RelationGetIndexedAttrs() (a per-call bms_copy) and the matching bms_free() still run for every index on every moved row. On the hot UPDATE path with several indexes this is avoidable palloc/free churn. Skip the copy when row_moved.
💡 Suggested change
Before:
indexedattrs = RelationGetIndexedAttrs(indexDesc);
indexInfo->ii_IndexNeedsUpdate =
row_moved || bms_overlap(indexedattrs, modified_idx_attrs);
bms_free(indexedattrs);
After:
if (row_moved)
indexInfo->ii_IndexNeedsUpdate = true;
else
{
indexedattrs = RelationGetIndexedAttrs(indexDesc);
indexInfo->ii_IndexNeedsUpdate =
bms_overlap(indexedattrs, modified_idx_attrs);
bms_free(indexedattrs);
}
📄 src/include/pgstat.h (L175-L176)
This new nontransactional counter is not reflected in the struct's header comment (line 163), which still enumerates tuples_inserted/updated/deleted/hot_updated/newpage_updated as the "count attempted actions" set. Add hot_indexed_updated to that list so the comment stays accurate with the field it now describes. (low confidence)
📄 src/backend/nodes/tidbitmap.c (L691-L692)
Resource-lifetime mismatch (moderate confidence). The VM buffer pinned here into a->vmcache is registered with whatever ResourceOwner is current during tbm_intersect() -- i.e. the ResourceOwner active in MultiExecBitmapAnd(). But the pin is only released in tbm_free(), which for the top-level bitmap runs in ExecEndBitmapHeapScan(), potentially under a different ResourceOwner (or after the pinning one has already been released at a subtransaction/portal boundary). When the pinning ResourceOwner is released first, it will emit a "buffer refcount leak" WARNING and drop the pin; the later ReleaseBuffer() in tbm_free() then operates on an already-released pin. Tie the pin to the bitmap's own lifetime deliberately (e.g. pin under a stable ResourceOwner, or release the cache at end of each tbm_intersect call rather than caching it in the long-lived TIDBitmap) rather than relying on tbm_free being the sole, correctly-scoped releaser.
📄 src/include/nodes/tidbitmap.h (L26-L27)
Pulling utils/relcache.h into this widely-included header widens the include graph for every consumer of tidbitmap.h just to name Relation in the new prototype. The existing convention in neighboring AM headers (amapi.h, amlocator.h) is to forward-declare the opaque struct instead. Prefer typedef struct RelationData *Relation; forward declaration (or struct RelationData * in the prototype) over the full relcache.h include.
💡 Suggested change
Before:
#include "utils/dsa.h"
#include "utils/relcache.h"
After:
#include "utils/dsa.h"
📄 src/include/access/amapi.h (L30-L32)
Unrelated whitespace-only churn: this removes a blank line in a file otherwise untouched by this feature. It serves no functional purpose and violates the minimal-diff discipline (it will also show up as noise in the pgsql-hackers patch). Drop this hunk.
📄 src/backend/utils/cache/relcache.c (L5430-L5444)
Memory-hygiene gap vs. the sibling helper in this same file. TextDatumGetCString(datum) palloc's a C string and stringToNode() builds a full parse tree, both in the caller's current memory context, and neither is ever freed. Compare RelationGetIndexExpressions() above, which does pfree(exprsString) right after stringToNode(). Here the string and the two parse trees (indexprs, indpred) leak. The leak is bounded to the first (uncached) call per relcache lifetime, but the first call can happen in a long-lived context (e.g. a query-lifetime context on the index-scan recheck path via heapam_index_entry_needs_recheck), so the leak persists. Free the string after parsing, and free the throwaway trees once pull_varattnos has consumed them.
💡 Suggested change
Before:
datum = heap_getattr(indexRel->rd_indextuple, Anum_pg_index_indexprs,
GetPgIndexDescriptor(), &isnull);
if (!isnull)
{
indexprs = (List *) stringToNode(TextDatumGetCString(datum));
pull_varattnos((Node *) indexprs, 1, &attrs);
}
datum = heap_getattr(indexRel->rd_indextuple, Anum_pg_index_indpred,
GetPgIndexDescriptor(), &isnull);
if (!isnull)
{
indpred = (List *) stringToNode(TextDatumGetCString(datum));
pull_varattnos((Node *) indpred, 1, &attrs);
}
After:
char *exprsString;
datum = heap_getattr(indexRel->rd_indextuple, Anum_pg_index_indexprs,
GetPgIndexDescriptor(), &isnull);
if (!isnull)
{
exprsString = TextDatumGetCString(datum);
indexprs = (List *) stringToNode(exprsString);
pfree(exprsString);
pull_varattnos((Node *) indexprs, 1, &attrs);
}
datum = heap_getattr(indexRel->rd_indextuple, Anum_pg_index_indpred,
GetPgIndexDescriptor(), &isnull);
if (!isnull)
{
exprsString = TextDatumGetCString(datum);
indpred = (List *) stringToNode(exprsString);
pfree(exprsString);
pull_varattnos((Node *) indpred, 1, &attrs);
}
📄 src/backend/utils/cache/relcache.c (L5382-L5393)
NULL-pointer dereference in the bootstrap fallback. indexStruct is indexRel->rd_index, and rd_index/rd_indextuple are always populated together (see RelationInitIndexAccessInfo line 1469-1470 and load_relcache_init_file line 6672; the init-file path even asserts both NULL together). So whenever rd_indextuple == NULL, rd_index is also NULL, and indexStruct->indnatts / indexStruct->indkey.values[i] dereference NULL. The guard protects against exactly the case its body then crashes on. Either this block is unreachable dead code (remove it) or it is reachable and must not touch indexStruct.
📄 src/backend/utils/cache/relcache.c (L5347-L5351)
This leading comment contradicts the actual implementation. The code deliberately does NOT use RelationGetIndexExpressions()/RelationGetIndexPredicate() (see the inline comment and the stringToNode() parsing of raw rd_indextuple columns below). It also states the copy is cached in rd_indexedattr, but the code caches it in rd_indattr (rd_indexedattr is the heap relation's field). Fix the comment to match the code.
📄 src/bin/pg_upgrade/relfilenumber.c (L596-L603)
High confidence: this VM-rewrite is wired only into transfer_relfile(), which --swap mode (TRANSFER_MODE_SWAP) bypasses entirely. transfer_single_new_db() dispatches swap mode to do_swap() before ever calling transfer_relfile(); do_swap()/swap_catalog_files() moves the old cluster's database directory verbatim and explicitly skips files matching the user-relation maps (bsearch continue), touching only catalog files. Consequently a 2-bit _vm fork from a pre-widening old cluster upgraded with --swap is moved into the new cluster unchanged and then misinterpreted under the 4-bit layout -- silent all-visible/all-frozen corruption and potentially wrong query results. The rewrite must also run on the swap-mode path (e.g. rewrite the _vm forks in place after the directory move, or forbid --swap when old cat_ver < VISIBILITYMAP_WIDTH_CHANGE_CAT_VER).
📄 src/bin/pg_upgrade/pg_upgrade.h (L127-L127)
Moderate confidence: VISIBILITYMAP_WIDTH_CHANGE_CAT_VER (202610071) currently equals CATALOG_VERSION_NO in this patch, but the comment itself admits this is a fragile hand-maintained magic constant that the committer must re-pick. catversion bumps are the committer's job at push time, and any clusters initdb'd with a catversion between an earlier HEAD and the final chosen value would still have 2-bit _vm forks yet compare as >= this threshold, so they would be copied verbatim and silently corrupted. Keying the rewrite decision on a self-maintained catversion constant is a footgun. Consider deriving the condition from an actual on-disk/control-file signal of VM width rather than a catversion comparison, or at minimum add a cross-check that fails loudly if the two drift.
📄 src/bin/pg_upgrade/t/009_hot_indexed.pl (L20-L22)
Moderate confidence: this test never exercises rewriteVisibilityMap(). Both old_node and new_node are initdb'd from the same HEAD build, so old_cluster.cat_ver >= VISIBILITYMAP_WIDTH_CHANGE_CAT_VER and the _vm rewrite branch in transfer_relfile() is never taken -- the _vm forks are copied verbatim. The most error-prone new code in this patch (the 2->4 bit VM transform, page roll-over flush/advance, final-page flush, and checksum stamping) therefore has zero coverage. Add a test that drives the rewrite path (e.g. a cross-version-style fixture or a mechanism to force the old cluster to look pre-widening) and verifies the resulting VM bits/checksums are correct.
📄 src/bin/pg_upgrade/file.c (L365-L365)
Low confidence: a partial read of a full page (read() returning 0 < n < BLCKSZ, which is permitted for regular files, e.g. when interrupted) exits the loop and is then reported as a fatal "partial page found", aborting an otherwise valid upgrade. The pre-existing copyFile() path in this same file loops on read() and tolerates short reads; consider using pg_pread() with a short-read retry loop (frontend convention) so a legitimately short read is completed rather than misclassified as a torn/partial page.
📄 src/bin/pg_upgrade/file.c (L338-L339)
Low confidence: this duplicates the on-disk VM layout arithmetic (MAPSIZE, HEAPBLOCKS_PER_PAGE, HEAPBLK_TO_) that lives in visibilitymap.c. The StaticAssertDecl only pins BITS_PER_HEAPBLOCK == 4; it does not guard the OLD_ 2-bit assumptions or MAPSIZE, so a future layout change to visibilitymap.c would silently diverge here and corrupt upgraded VMs. The 9.6 rewrite had the same hazard, but given this is format-critical, consider hoisting the shared packing macros into visibilitymapdefs.h so both sides stay in lockstep.
📄 contrib/amcheck/verify_heapam.c (L651-L654)
The stub detection relies solely on HotIndexedHeaderIsStub(), i.e. (t_infomask2 & HEAP_INDEXED_UPDATED) && natts == 0. For any LP_NORMAL item matching that loose predicate, all per-tuple validation (check_tuple) is skipped unconditionally. But the stub invariants this comment asserts -- HEAP_XMIN_INVALID | HEAP_XMAX_INVALID and HEAP_ONLY_TUPLE (see heap_page_prune's stub rewrite in pruneheap.c) -- are never actually verified here. A corrupt header that happens to have HEAP_INDEXED_UPDATED set and natts==0 but is otherwise damaged (live xmin, bogus xmax, not heap-only) will be silently classified as a legitimate stub and escape all corruption checks. Since amcheck's whole job is to detect such damage, validate the stub's own invariants (infomask == HEAP_XMIN_INVALID|HEAP_XMAX_INVALID, HEAP_ONLY_TUPLE set) and report_corruption when they do not hold, instead of blanket-skipping check_tuple on the mere presence of the two-bit signature.
📄 contrib/pageinspect/sql/hot_indexed_updates.sql (L25-L36)
get_hot_count is defined here and dropped at the end of the file, but it is never called anywhere in hot_indexed_updates.sql (every call site uses get_hi_count, which already exposes the hot column). This is dead scaffolding. Drop the definition (and its matching DROP FUNCTION get_hot_count(text) in the cleanup section) to keep the diff minimal, or add the missing call site if a classic-HOT-only check was intended here.
📄 contrib/pageinspect/sql/hot_indexed_updates.sql (L77-L82)
Comment/behavior mismatch and a plan-stability risk. This block only sets enable_seqscan = off; it does not disable bitmap scans, so the planner picks a Bitmap Heap Scan (confirmed in expected/hot_indexed_updates.out), not the nodeIndexscan path this comment describes. Two problems: (1) the comment is inaccurate about which executor node runs; (2) with COSTS OFF and only seqscan disabled, the IndexScan-vs-BitmapHeapScan choice is cost-based and can flip across planner/cost changes, making the pinned EXPLAIN output flaky. The later hi_range block correctly does SET enable_bitmapscan = off to force the IndexScan it intends to exercise; do the same here, or fix the comment to describe the Bitmap Heap Scan path actually tested.
📄 contrib/pageinspect/sql/hot_indexed_updates.sql (L860-L862)
Two distinct tables share the name hi_aba (here (k int, v int) with unique index hi_aba_k, and later (id int PRIMARY KEY, k int, v int) with a plain hi_aba_k). Each is dropped before the next is created, so this is currently valid, but the reused table and index names are a foot-gun: a future reorder, or an added test between them that forgets the intervening DROP, yields a confusing "relation already exists" failure. Consider renaming one (e.g. hi_aba_unique) to make the two cases unambiguous.
📄 src/bin/pg_upgrade/t/009_hot_indexed.pl (L27-L28)
This test does CREATE EXTENSION amcheck; and relies on verify_heapam() / bt_index_check(), but contrib/amcheck is not added to EXTRA_INSTALL in src/bin/pg_upgrade/Makefile (it only lists contrib/test_decoding src/test/modules/dummy_seclabel src/test/modules/test_extensions). Under the autoconf make check path amcheck will not be present in the temp install, so CREATE EXTENSION amcheck fails and the whole test errors out rather than being skipped. Add contrib/amcheck to EXTRA_INSTALL (and declare it as a dep in meson.build), or gate the amcheck-dependent assertions on availability. (high confidence)
📄 src/bin/pg_upgrade/pg_upgrade.h (L447-L450)
rewriteVisibilityMap() is defined in file.c (not relfilenumber.c), but its prototype is placed under the /* relfilenumber.c */ section. There is already a /* file.c */ section a few lines above (next to copyFile/linkFile); move this declaration there so the header's section comments stay accurate. (moderate confidence)
💡 Suggested change
Before:
/* relfilenumber.c */
void rewriteVisibilityMap(const char *fromfile, const char *tofile,
const char *nspname, const char *relname);
After:
/* relfilenumber.c */
📄 src/bin/pg_upgrade/file.c (L328-L329)
This rewrite is dispatched per 1GB segment by the segno loop in transfer_relfile(), but rewriteVisibilityMap() always restarts heapblk at 0 and writes all output into a single new_file. Widening from 2 to 4 bits doubles the map size, so a full old segment-0 _vm (up to RELSEG_SIZE) produces >1GB of output written to one file, overflowing the segment boundary; and for a multi-segment _vm (relfilenumber_vm.1, ...) each segment is transformed as if it started at heap block 0, giving wrong offsets. VM files large enough to span segments require an enormous heap (~32TB) so this is latent, but the segment math is incorrect. Consider documenting the single-segment assumption or handling segments explicitly. (low confidence)
📄 src/bin/pg_upgrade/file.c (L25-L28)
Include order deviates from the project convention of sorting headers alphabetically within a block: access/visibilitymapdefs.h should come before common/file_perm.h, and pg_upgrade.h (project header) is conventionally kept separate/after the sorted system+postgres headers. Reorder for consistency. (low confidence)
📄 contrib/amcheck/sql/check_heap.sql (L195-L197)
This scenario never creates a collapse-survivor stub, so the bulk of the new verify_heapam logic it claims to exercise (the pass-1 stub successor branch, the pass-2 stub-forwarding branch, and the stub arm of the redirect-intersection exemption) is left untested.
A stub is only recorded by prune when a dead mid-chain member changed an indexed attribute that is NOT changed again by a later hop (heap_prune_chain() -> heap_prune_record_stub(), gated on !HotIndexedBitmapIsSubset(attrs, laterattrs) in pruneheap.c). Here all three updates touch the same column c2, so each earlier hop's modified-attrs bitmap is a subset of the later hops' union and every dead member is reclaimed to LP_DEAD -- never stubbed. The chain collapses to redirects + LP_DEAD only.
To actually produce stubs, drive updates across different indexed columns so an earlier dead member's indexed change survives (is not superseded) when a later hop changes a different indexed column, e.g.:
UPDATE hot_indexed_check SET c1 = c1 + 1 WHERE id <= 50;
UPDATE hot_indexed_check SET c2 = c2 + 1 WHERE id <= 50;
UPDATE hot_indexed_check SET c1 = c1 + 1 WHERE id <= 50;
As written the test passes with 0 rows regardless of whether the stub paths are correct, so it would not catch a regression in them. (high confidence)
📄 contrib/pageinspect/sql/hot_updates.sql (L366-L373)
This test contradicts this file's own stated scope. The header (lines 5-11) says hot_updates.sql covers only HOT decisions that "apply identically on a pre-hot-indexed server: every UPDATE here either leaves all indexed attributes unchanged or touches only summarizing-index (BRIN) attributes", and that hot-indexed-specific behaviour (modifying a non-summarizing indexed attribute) "is covered in hot_indexed_updates.sql". This block updates tags, a column covered by a GIN index (non-summarizing), and both the comment ("a GIN-covered column changes, so this is HOT-indexed") and the expected output (hot = 1) assert the hot-indexed outcome. On a pre-hot-indexed server this UPDATE would be non-HOT (hot = 0), so the result is not server-variant-independent. Move this case to hot_indexed_updates.sql (which already has an equivalent GIN test, hi_gin) to keep the two files' scopes clean, or correct the file header. Leaving it here is a maintainability trap for the next reader.
📄 contrib/pageinspect/sql/hot_indexed_updates.sql (L57-L61)
hi_basic asserts exact chain-structure values from pg_relation_hot_indexed_stats (n_hot_indexed=1, n_chains=0, avg/max_chain_len=0 in the expected output) but does not set autovacuum_enabled = false. The later hi_reclaim / hi_vm tests deliberately set autovacuum_enabled = false and (for hi_reclaim) only assert horizon-independent facts (>= 1) precisely because opportunistic pruning / autovacuum can collapse the chain non-deterministically and change these exact counts. For consistency and to avoid a flaky exact-match assertion, add autovacuum_enabled = false to this table's storage parameters.
💡 Suggested change
Before:
+CREATE TABLE hi_basic (
+ id int PRIMARY KEY,
+ indexed_col int,
+ non_indexed_col text
+) WITH (fillfactor = 50);
After:
+CREATE TABLE hi_basic (
+ id int PRIMARY KEY,
+ indexed_col int,
+ non_indexed_col text
+) WITH (fillfactor = 50, autovacuum_enabled = false);
📄 contrib/pageinspect/sql/hot_indexed_updates.sql (L649-L650)
(hot - 0) is a no-op; hot - 0 is just hot. Write hot > 0 for clarity (minimal diff / readability nit).
💡 Suggested change
Before:
+SELECT (hot - 0) > 0 AS classic_hot_fired,
+ hot_idx = :hot_idx_before AS hot_indexed_did_not_fire
After:
+SELECT hot > 0 AS classic_hot_fired,
+ hot_idx = :hot_idx_before AS hot_indexed_did_not_fire
📄 src/test/benchmarks/siu/scripts/build.sh (L12-L18)
[high confidence] This benchmark harness is developer-local scaffolding, not a committable addition. It hardcodes absolute paths (/scratch/siu-bench, /scratch/pg, $HOME/ws/postgres/tepid) and bakes in a private branch name tepid that does not exist upstream. On any other machine or in CI these defaults resolve to nonexistent paths/revisions, so the scripts cannot run without a full set of env overrides. PostgreSQL has no src/test/benchmarks/ convention, and a performance claim belongs in the -hackers thread as a reproducible recipe, not as a committed per-developer tree. Recommend excluding this entire directory from the patch (minimal-diff discipline), or at minimum removing all /scratch and tepid-specific defaults.
📄 src/test/benchmarks/siu/scripts/split_ab.sh (L65-L65)
[medium confidence] Building split_off by sed -i.bak-rewriting a backend source file inside the working tree is fragile. The pattern matches the current source (.bitmap_and_inexact = heap_bitmap_and_inexact, exists verbatim in heapam_handler.c), but it is whitespace/format-sensitive: if pgindent reflows that line or the callback is renamed, sed silently matches nothing, split_off compiles identical to split_on, and the benchmark reports a bogus zero delta with no error. Worse, set -euo pipefail means any failure (compile error, signal) between the edit at this line and the mv ...c.bak restore leaves the backend source modified in the checkout. Add a post-sed assertion that the substitution actually changed the file, and restore the .bak from the EXIT trap rather than only on the success path.
📄 src/test/benchmarks/siu/scripts/run.sh (L363-L364)
[high confidence] The comment is wrong: UPDATE wide_table SET id=id is not a no-op. It still produces a new heap tuple (and, since id is the PK, a fresh index entry / HOT evaluation), so the wide_0 arm measures real write work, not a zero-column-change baseline. There is also no id % 1 anywhere here. Either pick a genuinely non-updating workload for the n=0 baseline or fix the comment to describe what actually happens.
📄 src/test/benchmarks/siu/scripts/hot_indexed_update.sql (L4-L6)
[high confidence] This workload references the pgbench variable :scale, but run.sh drives hot_indexed_update through the default *) branch, which passes no -D scale=... (only the wide_* and read_indexscan branches define their variables). pgbench aborts with "undefined variable scale" at \set aid, so this core A/B workload never actually runs; the harness records tps=NA for it. Either pass -D scale=$SCALE in the default branch or hardcode the row count the way read_indexscan does.
📄 src/test/benchmarks/siu/scripts/hot_indexed_mixed.sql (L3-L4)
[high confidence] Same :scale problem as hot_indexed_update.sql: this script is run via run.sh's default *) branch which does not pass -D scale=..., so pgbench fails on the undefined :scale variable and the hot_indexed_mixed workload produces no results.
📄 src/test/benchmarks/siu/scripts/build.sh (L12-L12)
[medium confidence] Default BENCH mismatch: the README documents the default bench root as /scratch/tepid-bench (and the build.sh example uses it), but every script actually defaults BENCH to /scratch/siu-bench (build.sh, run.sh, soak.sh, bloat.sh). Following the README verbatim points results/builds at a different directory than the scripts use. Pick one default and make the docs and all scripts agree (and drop the absolute /scratch path per the broader scaffolding concern).
📄 src/test/benchmarks/siu/scripts/run.sh (L266-L266)
[high confidence] Trailing whitespace after the closing quote. git diff --check flags this; strip it.
💡 Suggested change
Before:
local seedopt="--random-seed=$seed"
After:
local seedopt="--random-seed=$seed"
📄 src/test/regress/pg_regress.c (L1246-L1246)
This execl -> execlp change is unrelated to the HOT-indexed/SIU feature this patch implements, and it is a no-op for every supported build: shellprog is SHELLPROG, which both the Make build ($(SHELL), e.g. /bin/sh) and the meson build (hardcoded /bin/sh) define as an absolute path. When the program name contains a /, execlp skips the PATH search and behaves exactly like execl. The only situation where this matters is a custom environment that passes a bare shell name (see the added pg-aliases.sh/shell.nix dev tooling), so this looks like local dev-environment churn that leaked into the patch. Dropping it keeps the diff minimal; switching to a PATH-relative shell lookup for the test harness is also a mild footgun (behavior now depends on PATH) that would warrant its own justification.
💡 Suggested change
Before:
execlp(shellprog, shellprog, "-c", cmdline2, (char *) NULL);
After:
execl(shellprog, shellprog, "-c", cmdline2, (char *) NULL);
📄 src/test/regress/sql/tsearch.sql (L763-L764)
The reworded comment describes a feature-internal detail (the executor's SET-clause tracking introduced by this patch) rather than why this test exists. For the default heap build this test's behavior is unchanged, so if the comment is only meant to explain why the test still passes under the new in-place AM path, consider whether this core-test comment reword is needed at all; it is unrelated churn to tsearch and risks drifting from the implementation. If kept, make sure it describes behavior that has actually shipped and reads as commentary on the test, not on heapam internals.
📄 src/test/modules/test_locator_tableam/test_locator_tableam.c (L13-L13)
This comment line runs well past the tree's ~78-column wrapping limit, while every other line in this header block wraps at ~78. pgindent won't reflow block comments, so this will show up as an outlier. Re-wrap it to match the surrounding lines.
💡 Suggested change
Before:
* what such an AM would hand back. The old version is still on the page, so heap's own
After:
* what such an AM would hand back. The old version is still on the
* page, so heap's own
📄 src/test/regress/sql/triggers.sql (L663-L665)
This DELETE split (and the matching ORDER BY additions in generated_virtual.sql / updatable_views.sql) is a determinism fix: with a IN (20,21) two rows share a=21, so the INSTEAD OF DELETE trigger's per-row NOTICE order was unstable, and the split into unique (a,b) predicates makes it deterministic. The change itself is correct and the matching expected/*.out files are updated. Note for the committer: these core-test edits are NOT required for the default heap build (heap retains old row versions, so SIU does not reorder a seqscan); they only matter when the suite runs under an in-place table AM. That makes them scope-adjacent churn in otherwise unrelated core tests. If the intent is to run the standard regression suite under the new AM, that harness should be part of the patch so these edits are actually exercised; otherwise they are untested defensiveness.
📄 src/test/benchmarks/siu/scripts/soak.sh (L85-L87)
[high confidence] Two problems make this workload never run:
-
hot_indexed_update.sqlreferences the pgbench variable:scale, but this invocation passes no-D scale=...(unlike run.sh's read_indexscan/wide_* branches). pgbench aborts with "undefined variable scale". This is the same root cause as the run.sh:scalefinding, but soak.sh is a separate file with its own invocation that must set the variable, e.g.-D scale=$SCALE. -
soak.sh references
$BENCH/scripts/hot_indexed_update.sqlbut, unlike run.sh, never copies the scripts into$BENCH/scripts(it has no SRCDIR/cplogic). WhenBENCHis the default/scratch/siu-benchrather than this checkout, that path does not exist and pgbench fails with "could not open file". Either copy the scripts there first or reference them from the script's own directory.
💡 Suggested change
Before:
pgbench_as "$v" -f "$BENCH/scripts/hot_indexed_update.sql" \
-c "$CLIENTS" -j "$THREADS" -T "$DURATION" \
-P "$SAMPLE" -n postgres >"$LOGDIR/pgbench_$v.log" 2>&1 &
After:
pgbench_as "$v" -f "$BENCH/scripts/hot_indexed_update.sql" \
-c "$CLIENTS" -j "$THREADS" -T "$DURATION" \
-D "scale=$SCALE" -P "$SAMPLE" -n postgres >"$LOGDIR/pgbench_$v.log" 2>&1 &
📄 src/test/modules/test_locator_tableam/.gitignore (L1-L3)
This .gitignore omits /tmp_check/, which every other regress-based test module lists (e.g. test_tidstore, test_bloomfilter). Since this module has a C extension and a REGRESS target, make check spins up a temporary instance under tmp_check/; without this entry those generated files can be accidentally committed. Add it to match the tree-wide pattern.
💡 Suggested change
Before:
# Generated subdirectories
/log/
/results/
After:
# Generated subdirectories
/log/
/results/
/tmp_check/
📄 src/test/modules/test_locator_tableam/test_locator_tableam.c (L117-L117)
This offset check only guards the upper bound; it omits the >= FirstOffsetNumber lower-bound guard that the analogous core routine heap_get_latest_tid() performs (offnum < FirstOffsetNumber || offnum > PageGetMaxOffsetNumber(page)). If cur ever carries offset 0 (InvalidOffsetNumber) while following a t_ctid chain, PageGetItemId(page, 0) indexes before the first valid line pointer. Mirror the core guard for robustness.
💡 Suggested change
Before:
if (ItemPointerGetOffsetNumber(&cur) <= PageGetMaxOffsetNumber(page))
After:
if (ItemPointerGetOffsetNumber(&cur) >= FirstOffsetNumber &&
ItemPointerGetOffsetNumber(&cur) <= PageGetMaxOffsetNumber(page))
📄 src/test/regress/sql/updatable_views.sql (L128-L128)
Incomplete determinism fix. This hunk adds ORDER BY a only to the rw_view16 block (line 128), but the structurally identical rw_view15 block a few lines above (SELECT * FROM base_tbl; at line 121) is left unordered. At that point base_tbl holds 6 rows (the original -2..2 plus a=4, which came from UPDATE rw_view15 SET a=4 WHERE a=3), i.e. a multi-row SELECT immediately after an UPDATE -- the same pattern that motivated adding ORDER BY here. If the row-ordering change in this patch makes line 128 nondeterministic, line 121 is at least as exposed and will produce flaky expected/ output on the buildfarm/cfbot. Either both need ORDER BY or neither does; the asymmetry looks like a missed companion fix. (moderate confidence)
📄 src/test/modules/test_locator_tableam/sql/noretain.sql (L186-L186)
pg_ls_tmpdir() lists the entire default-tablespace pgsql_tmp directory cluster-wide, not just this session's temp files. These limit_kb and spilled assertions therefore assume no other backend (parallel worker, autovacuum, concurrent test) has temp files present at that instant. In particular, computing temp_file_limit as sum(size)/1024 + 40 and then expecting the following DO block to overshoot it (NOTICE: spill failed) is fragile: a stray temp file inflating sum(size) would raise the limit and the expected error would not fire, causing a spurious diff. Consider making the spill assertions tolerant of unrelated temp files, or document the reliance on an idle, isolated cluster.
📄 src/tools/pgindent/pgindent (L1-L1)
This shebang change is unrelated to the PR's stated purpose (the HOT-indexed access-method work) and should be split out. Beyond being an unrelated change, it is also inconsistent with the rest of the tree: every other Perl script in PostgreSQL (e.g. src/backend/catalog/genbki.pl, src/tools/copyright.pl, src/tools/mark_pgdllimport.pl, and ~30 others) uses #!/usr/bin/perl. Making pgindent the sole file using #!/usr/bin/env perl breaks that convention. The project deliberately standardizes on #!/usr/bin/perl, so this hunk should be dropped from the patch.
💡 Suggested change
Before:
#!/usr/bin/env perl
After:
#!/usr/bin/perl
📄 src/test/recovery/t/055_hot_indexed_recovery.pl (L88-L91)
This assertion does not verify that a prune/collapse happened, so the test is largely a worthless regression guard for the collapse path. n_hot_indexed counts live HOT-indexed versions (LP_NORMAL items with HEAP_INDEXED_UPDATED, natts>0). That flag is set by the HOT-indexed UPDATE itself (heapam.c sets HEAP_INDEXED_UPDATED on the new tuple), independent of pruning. So $pre_prune > 0 and $post_prune > 0 test the exact same invariant and both pass even if the opportunistic prune never ran and no chain ever collapsed to LP_REDIRECT/stubs. The metric that actually distinguishes "chain collapsed" from "chain never collapsed" is n_chains (LP_REDIRECT count) or max_chain_len, which this test never queries. Assert on those (e.g. n_chains > 0 / max_chain_len > 1 after the prune) so a revert of the collapse machinery would fail the test.
📄 src/test/recovery/t/055_hot_indexed_recovery.pl (L78-L79)
Factually wrong: heap_update() does not call heap_page_prune_opt(). The only heap_page_prune_opt() call site in heapam.c is in the sequential-scan path (heapgettup_pagemode), so it is the SELECT count(*) seqscan that drives the opportunistic prune here, not the subsequent UPDATE. heap_update only sets pd_prune_xid as a hint. Fix the comment to describe the real mechanism (scan-time prune_opt) and drop the "heap_update calls heap_page_prune_opt" claim. As written, the comment also admits uncertainty ("that's not enough on its own"), which signals the test's trigger is not actually understood.
📄 src/test/recovery/t/055_hot_indexed_recovery.pl (L138-L141)
The hardcoded "exactly two VACUUM passes" is fragile and risks flakiness. vacuumlazy.c's second heap pass explicitly only turns LP_DEAD into LP_UNUSED and "does NOT reclaim a collapsed HOT-indexed chain's stubs or re-point its redirects -- that chain-structure rewrite ... is left to a later prune." Whether that "later prune" happens within the second VACUUM, needs a third, or depends on the removal horizon/snapshot timing is implementation- and timing-dependent, not deterministically two. Across platforms and under the parallel schedule this is($final, '0', ...) can fail. Make reclamation deterministic (e.g. loop VACUUM until n_hot_indexed reaches 0 with a bounded retry, or assert via poll on the stat) rather than relying on a magic count of two.
📄 src/test/recovery/t/055_hot_indexed_recovery.pl (L2-L2)
Leading blank line before the copyright header. Standard PG test files start with the copyright line as line 1 (cf. 040_hot_indexed_replica_identity.pl and 009_hot_indexed.pl in this same change). Remove the blank first line for consistency.
📄 src/test/recovery/t/055_hot_indexed_recovery.pl (L4-L4)
This file collides with the existing t/055_cascade_reconnect.pl: recovery TAP tests use unique sequential numbers, and 056-059 are already taken. meson.build now lists two "055_" entries. Rename to the next free number (060_hot_indexed_recovery.pl) and update src/test/recovery/meson.build accordingly, otherwise the ordering/numbering convention is broken and the duplicate prefix is confusing. (high confidence)
📄 src/test/recovery/t/055_hot_indexed_recovery.pl (L122-L123)
Comment is factually wrong: verify_heapam's skip argument defaults to 'none' (see amcheck--1.2--1.3.sql: skip text default 'none'), not 'all-frozen'. The call correctly passes skip := 'none', but the comment contradicts it and misleads the reader. Fix the comment to say the default is 'none'. (high confidence)
💡 Suggested change
Before:
+# verify_heapam reports no errors on the relation (skip_option =
+# 'all-frozen' is the default; we want to scan everything).
After:
# verify_heapam reports no errors on the relation (skip := 'none' is the
# default; we pass it explicitly to scan everything).
📄 src/test/subscription/t/039_hot_indexed_apply.pl (L309-L310)
This subtest is vacuous: it does not exercise _bt_check_unique's stale-leaf recheck that the comment claims. The UNIQUE constraint is on (payload, tag). The seeded/updated row has key (999, 'mode_$mode') (live) with a stale leaf at (0, 'mode_$mode'). The INSERT uses tag 'fresh_$mode', so its key (0, 'fresh_$mode') differs on the tag component from both the live and the stale key -- there is no key collision at all, so _bt_check_unique never reaches the recheck path. This INSERT would succeed even if the staleness/recheck mechanism were completely broken, so the test passes with the feature reverted and catches no regression.
To actually exercise the recheck, the INSERT must collide with the stale leaf but not the live tuple -- i.e. reuse the same tag so the key matches the stale (0, 'mode_$mode') leaf while the live tuple is now (999, 'mode_$mode'):
💡 Suggested change
Before:
my ($r, $out, $err) = $subscriber->psql('postgres',
"INSERT INTO tab_uk VALUES ($ins_id, 0, 'fresh_$mode')");
After:
my ($r, $out, $err) = $subscriber->psql('postgres',
"INSERT INTO tab_uk VALUES ($ins_id, 0, 'mode_$mode')");
📄 src/test/subscription/t/040_hot_indexed_replica_identity.pl (L87-L88)
This test never confirms the HOT-indexed apply path actually fired on tab_full/tab_idx. If eligibility silently regressed to plain non-HOT (e.g. always-mode gating in HeapUpdateHotAllowable changed), the convergence checks and verify_heapam would still pass vacuously -- there would simply be no HOT-indexed chains/stubs to converge-around or to detect. 039 explicitly guards this (cmp_ok($ri_hotidx, '>', 0, ...) after polling n_tup_hot_indexed_upd); 040 should add the equivalent guard (poll n_tup_hot_indexed_upd on both tables and assert > 0) before the convergence/verify_heapam assertions.
📄 src/test/subscription/t/039_hot_indexed_apply.pl (L136-L137)
Flakiness risk: poll_counters only waits for the transactional counter n_tup_upd to reach its target, but the HOT-indexed/HOT counters it also returns are nontransactional (tab.counts_xact vs tab.trans->tuples_updated in pgstat_count_heap_update). They are flushed together by the same pgstat_report_stat, so this usually works, but there is no explicit target on n_tup_hot_indexed_upd/n_tup_hot_upd. For the cmp_ok(... '>', 0) assertions (Case 2 tab_pk HOT, Case 3 tab_extra HOT-indexed, tab_ri) a read taken right at the moment n_tup_upd crosses the target could in principle observe the subset of a flush. Consider polling the specific counter being asserted (e.g. pass the asserted counter and its target into the loop) rather than only n_tup_upd, to make the >0 assertions deterministic.
📄 src/test/subscription/t/039_hot_indexed_apply.pl (L2-L2)
Leading blank line before the copyright header. Standard PostgreSQL file headers place the copyright comment on line 1 (040 in this same patch does so, as do 037/038). Remove the empty first line so the header starts at line 1.
💡 Suggested change
Before:
# Copyright (c) 2026, PostgreSQL Global Development Group
After:
# Copyright (c) 2026, PostgreSQL Global Development Group
📄 src/test/subscription/t/039_hot_indexed_apply.pl (L324-L324)
Drop the personal reference from this committed-source comment. "Amit's corner" names an individual developer, which is noise to future readers and won't survive committer review. Describe the scenario instead, e.g. "Corner case:" or "Note:". Comments should explain the why, not credit a person.
💡 Suggested change
Before:
# Amit's corner: under hot_indexed_on_apply = 'always' the apply worker may
After:
# Corner case: under hot_indexed_on_apply = 'always' the apply worker may
| Diagnostics: | ||
| MissingIncludes: None |
There was a problem hiding this comment.
This file is personal developer-environment scaffolding and does not belong in a change destined for pgsql-hackers/commitfest. The repository's top-level .gitignore explicitly states: "Auxiliary files from local workflows, your preferred editor, etc. should be ignored locally using $GIT_DIR/info/exclude or ~/.gitexclude." A .clangd with a hardcoded build/ path and personal -DDEBUG -DLOCAL flags is exactly such a local-workflow file. Committing it violates the minimal-diff discipline and will be flagged as unrelated noise. Exclude it locally instead of committing it.
| # Main entry: heap_update | ||
| break heapam.c:3210 |
There was a problem hiding this comment.
This .gdbinit is personal developer tooling that should not be committed (see the repo's top-level .gitignore guidance on local-workflow files). More seriously, it is already stale/wrong against the very change it claims to annotate:
- The hardcoded source line numbers do not match the diff.
heap_update()begins near line 3296 in the new file, not 3210; the HOT update-decision block (if (hot_mode != HEAP_UPDATE_ALL_INDEXES)) is at ~line 4159, not 4019/4024/4033 (line 4019 is inside the TOAST insert path). Hardcoded absolute line numbers silently drift and will mislead anyone who sources this.
| # Predict augmented tuple size (returns 0 if t_hoff would overflow) | ||
| break heap_hot_indexed_tuple_size |
There was a problem hiding this comment.
These function-name breakpoints target symbols that do not exist anywhere in the change set. The real tuple-building function is heap_form_hot_indexed_tuple(), and the bitmap operations are inline helpers/macros in access/hot_indexed.h (e.g. HotIndexedSetAttrModified, HotIndexedBitmapUnion). heap_hot_indexed_tuple_size, heap_hot_indexed_create_tuple, and heap_hot_indexed_serialize_bitmap are not real symbols, so GDB will reject these break commands as unresolved.
| # Compute indexed_attrs for HOT indexed update chain following | ||
| break indexam.c:299 |
There was a problem hiding this comment.
This breakpoint targets indexam.c, which is not part of this change (indexam.c exists but is unmodified). The indexed_attrs computation referenced here lives in heapam_indexscan.c in this change, not indexam.c. The line number 299 is therefore meaningless for this patch.
| # WAL replay for XLOG_HEAP2_INDEXED_UPDATE | ||
| break heap_xlog_indexed_update |
There was a problem hiding this comment.
heap_xlog_indexed_update is not a real symbol in this change. WAL replay for updates goes through the existing heap_xlog_update path; the new XLH_UPDATE_NEW_LOCATOR_SPLIT flag is handled there, not in a dedicated heap_xlog_indexed_update function. This break will not resolve.
| @@ -0,0 +1,658 @@ | |||
| # PostgreSQL Development Aliases | |||
There was a problem hiding this comment.
This shell helper library is one developer's personal workstation setup and should not be committed to a PostgreSQL patch (see the repo's top-level .gitignore guidance on local-workflow files). Beyond the general minimal-diff concern, it hardcodes non-portable host assumptions: a trash command that is not standard, fixed /tmp/... paths, sudo tee /proc/sys/kernel/core_pattern sysctl manipulation, and a fixed regress port 40099 that will collide with any already-running instance on that port. None of this belongs in the tree.
| # Memory context debugging | ||
| define print_mcxt |
There was a problem hiding this comment.
These Nix dev-shell files are personal developer-environment scaffolding and do not belong in a change destined for pgsql-hackers/commitfest; they violate the minimal-diff discipline. If kept, two GDB helper macros are broken against current PostgreSQL structs and will error or print garbage when sourced, undercutting their stated debugging value:
print_mcxtreads$context->freepand$context->freeptr. The genericMemoryContextDatastruct (src/include/nodes/memnodes.h) has no such fields; those are AllocSet-era internals. GDB will fail to evaluate$context->freep - $context->freeptr.print_relcacheprints$rel->rd_rel->relnamespacewith%s, butrelnamespaceis anOid(pg_class.h), not a string, so this prints garbage or errors.
| inherit (shellConfig) gdbConfig flameGraphScript pgbenchScript; | ||
| }; | ||
|
|
||
| environment.localBinInPath = true; |
There was a problem hiding this comment.
environment.localBinInPath is a NixOS module option, not a valid flake output. Placed inside the per-system attrset returned by flake-utils.lib.eachDefaultSystem, it becomes a nonsensical top-level flake output (environment.<system>.localBinInPath) that no flake consumer or NixOS evaluation will ever read. It is dead configuration and should be removed (if the intent was to add the dev-shell install/bin to PATH, do it in the shellHook, not here).
| bool xs_prefix_all_dead; | ||
| } IndexScanHeapData; |
There was a problem hiding this comment.
The xs_prefix_all_dead field is written but never read. It is produced by heap_hot_search_buffer and stored into hscan->xs_prefix_all_dead (heapam_indexscan.c:762), reset at :766, but no consumer ever reads the struct field. The index-access layer that this comment says it "lets ... kill the arriving leaf" does not exist. (The local prefix_dead inside heap_hot_search_buffer is used internally, but the output field here is dead.) Either wire up the consumer or drop the field, its out-param, and the plumbing through heap_hot_search_buffer; speculative scaffolding and an aspirational comment both violate the minimalism/comment-accuracy rules.
| * require exact key agreement must skip such an entry; the row is | ||
| * re-supplied by the fresh entry inserted for the new value. Unlike | ||
| * xs_recheck (set by lossy AMs such as GiST and GIN), this verdict is | ||
| * produced by the table AM via table_index_entry_needs_recheck(); the |
There was a problem hiding this comment.
Stale/incorrect comment reference: table_index_entry_needs_recheck() does not exist anywhere in the tree. The verdict is produced by the static heapam_index_entry_needs_recheck() and written to this field directly in heapam_indexscan.c; there is no table-AM wrapper by this name. Fix the comment to name the real mechanism so it describes what the code does now.
| @@ -120,6 +120,7 @@ typedef struct ParallelBlockTableScanWorkerData | |||
| } ParallelBlockTableScanWorkerData; | |||
| typedef struct ParallelBlockTableScanWorkerData *ParallelBlockTableScanWorker; | |||
There was a problem hiding this comment.
Unrelated whitespace churn: this adds a second consecutive blank line before struct IndexScanInstrumentation; on an otherwise-untouched area. Drop it to keep the diff minimal (PostgreSQL patch hygiene; git diff --check/pgindent cleanliness).
| * The locator's size in bytes. A positive value is an exact size, the | ||
| * same for every row; a negative value is the largest a variable-width | ||
| * locator can be, so its magnitude bounds the width either way. | ||
| */ | ||
| int16 width; |
There was a problem hiding this comment.
The documented negative-width semantics ("a negative value is the largest a variable-width locator can be, so its magnitude bounds the width either way") is not implemented. The sole consumer, index_check_locator() in catalog/index.c, rejects the AM unconditionally when locdesc->width < 0 (it never interprets the magnitude). An AM declaring a variable-width locator would be refused outright regardless of its magnitude. Either implement the magnitude-bounds-the-width check in index.c or drop the negative-width documentation; describing a capability that does not exist is aspirational scaffolding.
| * leaves the top bit for the next user. A cluster upgraded from a 2-bit VM has | ||
| * its _vm forks rewritten by pg_upgrade. | ||
| */ | ||
| #define BITS_PER_HEAPBLOCK 4 |
There was a problem hiding this comment.
Widening BITS_PER_HEAPBLOCK from 2 to 4 silently breaks visibilitymap_count(). That function counts all-visible/all-frozen blocks via pg_popcount_masked() using VISIBLE_MASK8 (0x55) and FROZEN_MASK8 (0xaa) in visibilitymap.c (lines 135-136). Those masks are hard-coded for the old 2-bit/4-blocks-per-byte layout: 0x55 = 0101_0101 picks the low bit of each of four 2-bit pairs. With 4 bits/block (2 blocks/byte, bit0=visible, bit1=frozen, bit2=split), 0x55 now selects bit0 (visible, block A), bit2 (SPLIT, block A), bit4 (frozen, block B), bit6 (SPLIT, block B) -- so nvisible and nfrozen are garbage. This corrupts pg_class.relallvisible/relallfrozen and VACUUM's frozen accounting, misleading the planner. The masks must be updated for the 4-bit layout (visible -> 0x11, frozen -> 0x22). I grep-confirmed 0x55/0xaa appear nowhere else and are not updated by this patch. (high confidence)
| #define BITS_PER_HEAPBLOCK 4 | |
| #define BITS_PER_HEAPBLOCK 4 |
| /* Number of bits for one heap page */ | ||
| #define BITS_PER_HEAPBLOCK 2 | ||
| /* | ||
| * Bits per heap block in the visibility map. Widened from 2 to 4 to add |
There was a problem hiding this comment.
The pre-existing one-line comment "Number of bits for one heap page" is now stale and redundant: it sits directly above the new block comment that re-documents the same macro, and it says "page" where it means "block". For a minimal diff that reads as if always written this way, drop this leftover line and keep only the new block comment. (low confidence)
| * Bits per heap block in the visibility map. Widened from 2 to 4 to add | |
| * Bits per heap block in the visibility map. Widened from 2 to 4 to add |
| bool xs_index_only; /* caller is an index-only scan that may | ||
| * return tuples without fetching the heap; | ||
| * AMs must retain leaf-page pins for such | ||
| * scans (VM all-visible / TID-recycle race), | ||
| * whereas a plain scan that sets xs_want_itup | ||
| * only to inspect the index tuple still | ||
| * fetches the heap and may drop pins */ |
There was a problem hiding this comment.
xs_index_only is documented as the flag that forces AMs to retain leaf-page pins for genuine index-only scans, but it is never assigned true anywhere in the tree. It is initialized to false in genam.c ("may be set later") and consumed in nbtinsert/nbtree only: btrescan() changed so->dropPin from !scan->xs_want_itup to !scan->xs_index_only. Because index_only_scan still sets only xs_want_itup = index_only_scan (indexam.c index_beginscan_internal), an index-only scan now has xs_index_only == false, so with an MVCC snapshot and a heap relation dropPin becomes true and nbtree drops leaf-page pins during index-only scans. That reintroduces exactly the VM all-visible / TID-recycle race the field's comment says must be prevented -- a data-correctness bug (wrong results, worst on hot standby). Set scan->xs_index_only = true for index-only scans in index_beginscan_internal() alongside xs_want_itup = index_only_scan.
| so->dropPin = (!scan->xs_index_only && | ||
| IsMVCCLikeSnapshot(scan->xs_snapshot) && | ||
| scan->heapRelation != NULL); |
There was a problem hiding this comment.
This re-keys dropPin on xs_index_only, but nothing in the tree ever sets scan->xs_index_only = true: it is only ever initialized to false in RelationGetIndexScan() (genam.c), and index_beginscan_internal() sets xs_want_itup = index_only_scan but not xs_index_only. As a result !scan->xs_index_only is always true for an MVCC scan with a heap relation, so dropPin becomes true for genuine index-only scans too. That is exactly the VM all-visible / TID-recycle race the preceding comment (and the nbtree README) says must be avoided -- index-only scans will now drop the leaf pin while relying on the VM. The feature must set xs_index_only = true on the index-only-scan path (e.g. in index_beginscan_internal alongside xs_want_itup) before this change is safe. Confidence: high.
| scan->orderByData = NULL; | ||
|
|
||
| scan->xs_want_itup = false; /* may be set later */ | ||
| scan->xs_index_only = false; /* may be set later */ |
There was a problem hiding this comment.
xs_index_only is introduced here (and relied upon by nbtree's dropPin logic) but is never set to true anywhere in the tree -- index_beginscan_internal() sets xs_want_itup = index_only_scan without a corresponding xs_index_only = index_only_scan. The "may be set later" comment is therefore inaccurate: no producer exists, so index-only scans fall into the pin-dropping path and reintroduce the documented VM all-visible / TID-recycle race. Wire up the setter on the index-only-scan path. Confidence: high.
| /* ------------------------------------------------------------------------ | ||
| * Functions for non-modifying operations on individual tuples |
There was a problem hiding this comment.
Unrelated whitespace churn: this adds a third blank line before the "non-modifying operations" banner, which is not required by the signature changes in this file. PostgreSQL rejects reformatting of untouched lines; drop this hunk to keep the diff minimal and pgindent-clean. (low confidence, style)
| /* ------------------------------------------------------------------------ | |
| * Functions for non-modifying operations on individual tuples | |
| /* ------------------------------------------------------------------------ | |
| * Functions for non-modifying operations on individual tuples |
| /* ------------------------------------------------------------------------ | ||
| * Functions for non-modifying operations on individual tuples |
There was a problem hiding this comment.
Unrelated whitespace churn: two extra blank lines inserted after table_index_getnext_slot(), unrelated to this change. Drop to keep the diff minimal and pgindent-clean. (low confidence, style)
| /* ------------------------------------------------------------------------ | |
| * Functions for non-modifying operations on individual tuples | |
| /* ------------------------------------------------------------------------ | |
| * Functions for non-modifying operations on individual tuples |
| bool (*fetch_tid_check) (Relation rel, | ||
| ItemPointer tid, | ||
| Snapshot snapshot, | ||
| bool *all_dead, | ||
| bool *recheck, | ||
| TupleTableSlot *slot); |
There was a problem hiding this comment.
The continuation parameter lines of fetch_tid_check are indented with a flat offset instead of aligning under the opening paren the way every sibling callback here does (compare tuple_update/modified_attrs/tuple_lock just below). This will not pass pgindent cleanly; align the parameters. (low confidence, style)
| TupleTableSlot *slot); | ||
|
|
||
|
|
||
| /* ------------------------------------------------------------------------ | ||
| * Callbacks for non-modifying operations on individual tuples |
There was a problem hiding this comment.
Extra blank line added after the fetch_tid_check callback (two blank lines before the "non-modifying operations" banner). Unrelated reformatting; keep a single blank line to match the surrounding style and keep the diff minimal. (low confidence, style)
| rel->rd_locdesc = rel->rd_tableam->relation_locator ? | ||
| rel->rd_tableam->relation_locator(rel) : &default_locdesc; |
There was a problem hiding this comment.
Footgun: RelationGetLocatorDesc() dereferences rel->rd_tableam->relation_locator without a NULL check on rd_tableam, yet the sibling RelationUpdatesInPlace() just below explicitly guards rel->rd_tableam != NULL and documents "False for a relation without a table AM, such as a foreign table." This asymmetry means a direct call to RelationGetLocatorDesc() on a relation without a table AM (view, foreign table, index) will crash on a NULL deref, while the comment on this function only mentions AMs that "predate the locator contract" (i.e. rd_tableam set but callback NULL), not rd_tableam == NULL. Current in-tree callers all pass heap/table relations so this is not a live bug, but either guard rd_tableam here too or document that callers must only pass relations that have a table AM. (medium confidence, maintainability)
| /* | ||
| * A redirect target (LP_REDIRECT) is a valid chain root: an index entry | ||
| * pointing at it is legitimate and the caller's chain walk decides | ||
| * deletability. Only genuinely normal tuples are inspected below. |
There was a problem hiding this comment.
Loss of corruption detection for non-SIU relations. The previous code reported ERRCODE_INDEX_CORRUPTED when an index entry pointed at a heap-only tuple (if (unlikely(HeapTupleHeaderIsHeapOnly(htup)))). This block replaces that check with a comment and performs no check at all. The two earlier checks in this function were correctly gated on inexact_ok (keeping the ereport for ordinary relations), but this heap-only check was removed unconditionally, so a corrupt index entry pointing at a heap-only tuple on a plain (non-SIU) table is now silently accepted instead of detected. Gate the tolerance on inexact_ok and keep the ereport for !inexact_ok.
|
📊 OCR posted 23/25 inline comment(s). 2 could not be posted
|
No description provided.