<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Nobody]]></title><description><![CDATA[Nobody]]></description><link>https://getnobody.hashnode.dev</link><image><url>https://cdn.hashnode.com/res/hashnode/image/upload/v1593680282896/kNC7E8IR4.png</url><title>Nobody</title><link>https://getnobody.hashnode.dev</link></image><generator>RSS for Node</generator><lastBuildDate>Tue, 29 Sep 2026 06:46:42 GMT</lastBuildDate><atom:link href="https://getnobody.hashnode.dev/rss.xml" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><ttl>60</ttl><item><title><![CDATA[Why your SQLite WAL file never shrinks]]></title><description><![CDATA[Originally published at https://getnobody.org/blog/sqlite-wal-never-shrinks/
If you run SQLite in WAL mode with a busy writer and a few readers, you may one day find a -wal file of several gigabytes n]]></description><link>https://getnobody.hashnode.dev/why-your-sqlite-wal-file-never-shrinks</link><guid isPermaLink="true">https://getnobody.hashnode.dev/why-your-sqlite-wal-file-never-shrinks</guid><dc:creator><![CDATA[Nobody]]></dc:creator><pubDate>Fri, 25 Sep 2026 15:19:17 GMT</pubDate><content:encoded><![CDATA[<p>Originally published at <a href="https://getnobody.org/blog/sqlite-wal-never-shrinks/">https://getnobody.org/blog/sqlite-wal-never-shrinks/</a></p>
<p>If you run SQLite in WAL mode with a busy writer and a few readers, you may one day find a -wal file of several gigabytes next to a database that is much smaller. Autocheckpoint is on, nothing reports an error, and PRAGMA wal_checkpoint says it worked. I read the checkpoint code in SQLite 3.53.2, then built a small lab to watch it happen.</p>
<p>The short version: when read transactions overlap so that at least one is always open, the checkpoint can never finish and the writer can never go back to frame 1. In my runs the WAL grew by 11.9 MiB per second for as long as the overlap lasted, and when the readers went away it stayed at 357 MiB. A wal_checkpoint(TRUNCATE) every 5 seconds kept it under 60 MiB in 2 of 3 runs (119 MiB in the run where one call came back busy), at the price of stalling the writer for up to 1.8 s each time.</p>
<h2>How WAL mode commits and checkpoints</h2>
<p>In WAL mode a commit doesn't touch the database file. The writer appends the changed pages to the -wal file as frames (a 24-byte header plus one page), and a shared-memory index in the -shm file maps page numbers to the newest frame that holds them. Readers look there first and fall back to the database file.</p>
<p>A checkpoint copies frames back into the database file. Two counters in the shared index drive everything (wal.c L321-L401):</p>
<p><code>mxFrame</code> is the last committed frame. <code>nBackfill</code> is how far the checkpoint has copied. When <code>nBackfill</code> reaches <code>mxFrame</code>, everything in the WAL is in the database, and the next writer can start again at frame 1 instead of appending. It overwrites the old file from the start. It doesn't truncate it.</p>
<p>By default SQLite runs a checkpoint after any commit that leaves the WAL at 1,000 frames or more (sqlite3WalDefaultHook, main.c L2470-L2481). That checkpoint is the PASSIVE kind: it never waits for anybody.</p>
<p>(Chart: see the original post.)</p>
<h2>What a reader pins</h2>
<p>A reader doesn't copy anything. When it starts a read transaction it pins one of four read marks (<code>aReadMark[1..4]</code>, since WAL_NREADER is 5 and slot 0 is special) whose value is no greater than the <code>mxFrame</code> it saw. If a slot already holds that value it shares the slot with the readers already on it. Otherwise it writes its <code>mxFrame</code> into a slot it can lock exclusively, and if it can't get one, it settles for the largest older mark (walTryBeginRead, L3163-L3199). It holds a shared lock on that slot until the transaction ends.</p>
<p>Frames after its mark are newer than its snapshot. For pages changed in those frames, the reader still needs the old copy in the database file, so the checkpoint may not copy those frames over it until the reader finishes (L378-L381). Frames up to its mark are safe to copy.</p>
<p>If the WAL is fully copied when the reader starts, it takes slot 0 instead, which means "read the database file only, ignore the WAL" (walTryBeginRead, L3128-L3160). Remember that slot 0, it matters later.</p>
<h2>Why the checkpoint stops at the oldest reader</h2>
<p>walCheckpoint starts by assuming it can copy everything, then walks the four read slots (L2227-L2245):</p>
<pre><code class="language-c">mxSafeFrame = pWal-&gt;hdr.mxFrame;
mxPage = pWal-&gt;hdr.nPage;

for(i=1; i&lt;WAL_NREADER; i++){
  u32 y = AtomicLoad(pInfo-&gt;aReadMark+i);
  SEH_INJECT_FAULT;

  if( mxSafeFrame&gt;y ){
    assert( y&lt;=pWal-&gt;hdr.mxFrame );

    rc = walBusyLock(
      pWal,
      xBusy,
      pBusyArg,
      WAL_READ_LOCK(i),
      1
    );

    if( rc==SQLITE_OK ){
      u32 iMark = (i==1 ? mxSafeFrame : READMARK_NOT_USED);

      AtomicStore(pInfo-&gt;aReadMark+i, iMark);
      SEH_INJECT_FAULT;
      walUnlockExclusive(pWal, WAL_READ_LOCK(i), 1);

    }else if( rc==SQLITE_BUSY ){
      mxSafeFrame = y;
      xBusy = 0;

    }else{
      goto walcheckpoint_out;
    }
  }
}
</code></pre>
<p>If it can lock a slot exclusively, nobody is using that mark, so it resets it: slot 1 to <code>mxSafeFrame</code>, the others to <code>READMARK_NOT_USED</code> (<code>0xffffffff</code>, unused). If it can't, a live reader is there, and <code>mxSafeFrame</code> drops to that reader's mark. The copy loop then skips every frame after <code>mxSafeFrame</code> (L2306), and <code>nBackfill</code> only advances to it (L2331).</p>
<p>Then comes the line that makes this hard to notice (L2340-L2345):</p>
<pre><code class="language-c">if( rc==SQLITE_BUSY ){
  /* Reset the return code so as not to report a checkpoint failure
  ** just because there are active readers. */
  rc = SQLITE_OK;
}
</code></pre>
<p>So PRAGMA wal_checkpoint(PASSIVE) returns busy = 0 even when it copied only part of the WAL. The signal is in the other two columns: log (frames in the WAL) keeps getting bigger than checkpointed.</p>
<h2>Why the file grows, and why it stays big</h2>
<p>Copying is only half of it. The WAL goes back to frame 1 in walRestartLog (L3869-L3909), which runs at the start of every write. It rewinds only if the writer's own snapshot is on slot 0 (<code>pWal-&gt;readLock==0</code>, so the WAL was fully copied when the writer started), and it can lock read slots 1 to 4 exclusively, meaning no reader is using the WAL.</p>
<p>Now put a steady writer next to readers that overlap. Some reader always holds a mark from a second or so ago. The checkpoint copies up to that mark and stops. <code>nBackfill</code> never catches up with <code>mxFrame</code>, so the writer is never on slot 0, so it never rewinds. Every commit appends. The -shm index grows with it, one 32 KiB block per 4,096 frames (HASHTABLE_NPAGE, L615).</p>
<p>The SQLite docs describe exactly this under checkpoint starvation: "if a database has many concurrent overlapping readers and there is always at least one active reader, then no checkpoints will be able to complete and hence the WAL file will grow without bound."</p>
<p>And when the overlap ends, the file keeps its size. A rewind moves the write position to frame 1, but, as the same page says, a checkpoint "does not normally truncate the WAL file (unless the journal_size_limit pragma is set)". The file only gets smaller when:</p>
<ul>
<li><p>journal_size_limit is set, and the first commit after a rewind trims the file to that limit (L4210-L4221). The default in this build is -1, no limit.</p>
</li>
<li><p>someone runs wal_checkpoint(TRUNCATE), which truncates it to 0 bytes (L2363-L2381).</p>
</li>
<li><p>the last connection to the database closes, checkpoints and deletes the WAL.</p>
</li>
</ul>
<h2>The experiment</h2>
<p>One writer process commits one row with a 1,000-byte blob per transaction, open loop at a fixed 2,000 commits per second, with synchronous=NORMAL. Reader processes each run BEGIN, a SELECT that takes the snapshot, hold it, look up 100 recent rows, COMMIT, then idle. A sampler records the file sizes every 100 ms.</p>
<p>Seven scenarios, 30 s each, 3 runs of each, interleaved:</p>
<ul>
<li><p>Scenario a writer only</p>
</li>
<li><p>b 2 readers holding 0.4 s every 1.0 s</p>
</li>
<li><p>c 3 readers holding 1.0 s, idling 0.2 s, staggered: at least 2 are always open</p>
</li>
<li><p>d c, plus wal_checkpoint(TRUNCATE) every 5 s with a 5 s busy timeout</p>
</li>
<li><p>e c, plus journal_size_limit of 4 MiB</p>
</li>
<li><p>f b, plus journal_size_limit of 4 MiB</p>
</li>
<li><p>g c, with wal_autocheckpoint=0 on the writer (the sampler's once-a-second probe still checkpoints, see below)</p>
</li>
</ul>
<p>Here is what the readers actually did in 3 seconds of one run of b and one of c, straight from the logged open and close times:</p>
<p>(Chart: see the original post.)</p>
<p>And here is the WAL file:</p>
<p>(Chart: see the original post.)</p>
<p>With gaps (b), the WAL rewinds twice a second, once per gap, 60 times in 30 s, and peaks at 6.05 to 6.08 MiB. With overlap (c), it never rewinds once, and grows in a straight line at 11.91 MiB/s in all three runs. That is about 1.5 frames of 4,120 bytes per commit: the final checkpoint counted 90,873 frames for 60,000 commits. At the end the -shm file was 736 KiB, 23 index blocks.</p>
<p>Through all of c, every PASSIVE probe returned busy = 0, while log ran up to 2,751 frames ahead of checkpointed. That is a little under one second of writes, about one reader hold time.</p>
<h2>After the readers leave</h2>
<p>When the run ended I stopped the readers and ran one more PASSIVE checkpoint with nothing else open. It copied everything, returned <code>0 | 90873 | 90873</code>, and the file stayed where it was:</p>
<p>(Chart: see the original post.)</p>
<p><code>journal_size_limit</code> did nothing in e. Its trim runs on the first commit after a rewind, and under overlap there are no rewinds. In f, where there are, it held the file at 4.00 MiB, though the high-water mark still reached 6.01 MiB between trims. In one of the three e runs the file was already 4.00 MiB before that final checkpoint. Most likely the readers closed a moment before the writer's last commits during shutdown, so one commit found a fully copied WAL, rewound and was trimmed: the same mechanism, with the blocker gone.</p>
<h2>What TRUNCATE costs</h2>
<p>The non-passive modes behave differently (sqlite3WalCheckpoint, L4295-L4426). FULL, RESTART and TRUNCATE take the writer lock first, waiting with the busy handler. Because xBusy is set, walBusyLock then waits on each reader slot instead of giving up. RESTART and TRUNCATE also wait until no reader is using the WAL at all (L2358-L2362), and TRUNCATE then cuts the file to 0 bytes. Readers that start after the WAL is fully copied take slot 0, which doesn't hold the checkpoint up, and in my runs the readers' 100-row lookup was no slower in d than in c (226 against 224 µs median). The SQLite docs do warn that readers might block while it runs. The writer is blocked for sure, because the checkpoint holds its lock.</p>
<p>In d, 14 of 15 TRUNCATE calls reset the WAL and returned <code>0 | 0 | 0</code>. They took 1.40 to 1.81 s (median 1.70 s) with readers that hold for 1.0 s. That is close to two hold times, and the code explains why: first the call waits for the readers pinned below mxFrame, but while it waits, readers that start see a WAL that isn't fully copied yet, so they take a real mark at the new mxFrame. The RESTART step then has to wait for those too. On all 14 calls, the checkpoint returned within 105 ms of the last close of a reader that had started before the copy finished.</p>
<h2>Results</h2>
<table>
<thead>
<tr>
<th></th>
<th>Median of 3 runs c: overlap</th>
<th>d: c + TRUNCATE every 5 s</th>
</tr>
</thead>
<tbody><tr>
<td>WAL at 30 s</td>
<td>357.03 MiB</td>
<td>59.58 MiB</td>
</tr>
<tr>
<td>Writer time inside execute(), p50 / p99.9</td>
<td>38 / 2,653 µs</td>
<td>38 / 745 µs</td>
</tr>
<tr>
<td>Writer response time p99</td>
<td>0.50 ms</td>
<td>1,719 ms</td>
</tr>
<tr>
<td>Writer response time max</td>
<td>6.7 ms</td>
<td>1,814 ms</td>
</tr>
<tr>
<td>Commits more than 100 ms late (of 50,000)</td>
<td>0</td>
<td>16,442</td>
</tr>
</tbody></table>
<p>"Response time" is when a commit finished minus when the open-loop schedule said it should start, so a stall counts in full. Time inside execute() hides it: only one commit per TRUNCATE actually waits, and the thousands queued behind it run fast once it lets go. If you measure the effect of a checkpoint job on your writer, measure response time.</p>
<p>The one TRUNCATE that failed came back in 1.68 ms with <code>1 | -1 | -1</code>. That is the WAL_CKPT_LOCK check at the very start (L4336): if another checkpoint is already running, every mode returns busy without calling the busy handler. My probe wasn't running at that moment, so the other checkpoint was almost certainly the writer's own autocheckpoint, which under starvation runs after every commit. That call missed, the WAL grew for another 5 s, and that run's high-water mark was 118.86 MiB instead of about 60.</p>
<h2>Two smaller results</h2>
<p>Checkpointing on the writer's connection costs it something under starvation. Once the WAL is past 1,000 frames, the autocheckpoint runs after every commit and copies whatever it can. In g the writer's autocheckpoint was off, and the only checkpoint was the sampler's PASSIVE probe once a second from another connection. That dropped the writer's p50 from 38 to 32 µs and its p99.9 from 2,653 to 694 µs. The WAL grew exactly as fast, and the probe kept nBackfill as close as the autocheckpoint did (at most 2,743 frames behind, against 2,751 in c), since the checkpoint wasn't the thing holding the WAL down anyway. Reads did not get slower, and the code says they shouldn't. A reader doesn't search the whole WAL index. walFindFrame walks hash blocks only from its mxFrame down to the block holding minFrame (L3568-L3569), and minFrame is nBackfill + 1 when the read transaction starts (L3253). So lookup cost depends on how far the checkpoint lags, not on the size of the WAL. Under starvation the lag stayed at most 2,751 frames, less than one 4,096-frame block, so a lookup searched 1 or 2 blocks. In c the median was 216 µs for transactions that started 5 to 12 s into the run and 218 µs for 22 to 30 s, at a WAL of up to 357 MiB. Reads would slow down only if nBackfill itself fell far behind, for example with autocheckpoint off and no other checkpointer. I didn't measure that.</p>
<h2>What to do in your app</h2>
<p>Keep read transactions short. The problem is overlap, and overlap comes from long snapshots: a cursor left open while you call an API, a BEGIN that a connection pool never closed. In Python, a SELECT whose rows you haven't fully fetched keeps its statement active, and an active statement keeps its read snapshot. Watch for it.</p>
<p>Poll PRAGMA wal_checkpoint(PASSIVE) and alert when log - checkpointed keeps rising, and watch the size of the -wal file. The busy column won't tell you.</p>
<p>Set journal_size_limit, for example:</p>
<pre><code class="language-sql">PRAGMA journal_size_limit = 67108864
</code></pre>
<p>for 64 MiB. It doesn't stop starvation, but it gives back the disk the next time the WAL rewinds.</p>
<p>Run wal_checkpoint(TRUNCATE) or (RESTART) from a maintenance job with a busy timeout, when writes are quiet. Expect it to block the writer for up to about twice your longest read transaction. If it returns busy = 1, retry soon rather than waiting for the next period.</p>
<p>If you turn off autocheckpoint, checkpoint from another connection. In g, one PASSIVE checkpoint a second from a separate connection kept up as well as the autocheckpoint and took that work off the writer. With nothing checkpointing at all, nBackfill never reaches mxFrame, so the WAL never rewinds, gaps or not.</p>
<h2>How I measured</h2>
<p>Machine: Apple M5 Pro (18 cores, 48 GiB), macOS 26.6.2, APFS on the internal SSD.</p>
<p>Software: Python 3.14.4 and the SQLite it links, 3.53.2 (DEFAULT_WAL_AUTOCHECKPOINT=1000, DEFAULT_JOURNAL_SIZE_LIMIT=-1, from PRAGMA compile_options).</p>
<p>The source links above are pinned to the version-3.53.2 tag, because that is what Python links here. 3.53.3 and 3.53.4 are out.</p>
<p>Runs: 7 scenarios × 3 runs × 30 s, interleaved. The first 5 s of each run are left out of the latency statistics. Tables show the median of 3 runs.</p>
<p>The writer is open loop, so every run issued all 60,000 scheduled commits, and all of them completed with no write errors. That says nothing about stalls; the response times above do.</p>
<p>Caveats: The sampler's PASSIVE probe once a second is itself a checkpoint, and it runs in every scenario. Where autocheckpoint is on, it copies no more than the autocheckpoint would. In g it is the only checkpointer, so g means "one PASSIVE checkpoint a second from another connection", not "no checkpointing". --probe-every 0 turns it off.</p>
<p>The first full pass had a probe colliding with the TRUNCATE calls and a maintenance call firing after the run had ended. I fixed both and reran everything. All numbers here are from the second pass.</p>
<p>This is one machine and one filesystem. I did not run the Linux cross-check or a sweep of write rates. "MiB" is 2^20 bytes.</p>
<p>The scripts, the raw CSVs and how to rerun them are in the lab folder. <code>python3 run.py</code> runs the full matrix in about 12 minutes with nothing but the standard library.</p>
<h2>Credit</h2>
<p>Credit goes to D. Richard Hipp and the SQLite developers, who wrote and maintain the WAL code and whose WAL documentation already describes this failure and the "reader gaps" fix in plain words. The code is commented well enough that a stranger can follow a checkpoint line by line, which is how this post got written.</p>
<p>The same page also documents this year's WAL-reset bug, a rare corruption case when two connections write or checkpoint at the same instant, which one of the SQLite developers, Dan, found and fixed in 3.51.3.</p>
]]></content:encoded></item><item><title><![CDATA[How Postgres 19 checks foreign keys without running SQL]]></title><description><![CDATA[Every time you insert a row into a table with a foreign key, Postgres has to prove the referenced row exists and stop anyone from deleting it until your transaction ends. Up to Postgres 18, it did tha]]></description><link>https://getnobody.hashnode.dev/how-postgres-19-checks-foreign-keys-without-running-sql</link><guid isPermaLink="true">https://getnobody.hashnode.dev/how-postgres-19-checks-foreign-keys-without-running-sql</guid><category><![CDATA[PostgreSQL]]></category><category><![CDATA[Databases]]></category><category><![CDATA[performance]]></category><dc:creator><![CDATA[Nobody]]></dc:creator><pubDate>Thu, 24 Sep 2026 23:53:41 GMT</pubDate><content:encoded><![CDATA[<p>Every time you insert a row into a table with a foreign key, Postgres has to prove the referenced row exists and stop anyone from deleting it until your transaction ends. Up to Postgres 18, it did that by running a tiny SQL query for every single row. Postgres 19 stops doing that. I read the new code, then measured it.</p>
<p>The short version: on a 1 million row bulk insert, the foreign key check costs about <strong>44% less time</strong> in Postgres 19. A second optimization that was pulled from 19 before release would have helped much more, but only when the foreign key values arrive in order.</p>
<h2>The old path: a query per row</h2>
<p>A foreign key is enforced by an internal <code>AFTER INSERT</code> trigger, <code>RI_FKey_check</code> in <a href="https://github.com/postgres/postgres/blob/8fbbb338bed77c6fa8dc6fa309cbcca93c1c2d72/src/backend/utils/adt/ri_triggers.c"><code>ri_triggers.c</code></a>. In Postgres 18 that trigger builds this query once, saves the plan, and runs it through SPI (the server's internal SQL interface) for every inserted row:</p>
<pre><code class="language-sql">SELECT 1 FROM ONLY parent x WHERE id = $1 FOR KEY SHARE OF x
</code></pre>
<p><code>FOR KEY SHARE</code> is the important part. It locks the parent row so nobody can delete it or change its key while your transaction is open, but it still lets other sessions update the parent's non-key columns.</p>
<p>The plan is cached, but every row still pays for SPI setup, executor startup and teardown, snapshot handling and the permission checks, just to do one index lookup.</p>
<img src="https://cdn-images-1.medium.com/v2/resize:fit:1600/1*Yrwz6WZOFhIAYW2vcihpgw.png" alt="Postgres 18 checks each foreign key through SPI, a saved plan and the executor. Postgres 19 calls ri_FastPathCheck, which probes the btree index and locks the row directly." style="display:block;margin:0 auto" />

<p><em>Same check, same lock. Postgres 19 removes everything between the trigger and the index.</em></p>
<h2>The new path: probe the index directly</h2>
<p>Amit Langote's <a href="https://github.com/postgres/postgres/commit/2da86c1ef9b5446e0e22c0b6a5846293e58d98e3">commit 2da86c1ef9</a> adds a fast path. Before touching SPI, <code>RI_FKey_check</code> now asks two questions:</p>
<pre><code class="language-c">if (ri_fastpath_is_applicable(riinfo) &amp;&amp;
    ri_FastPathCheck(riinfo, fk_rel, newslot))
    return PointerGetDatum(NULL);
</code></pre>
<p><code>ri_fastpath_is_applicable</code> says no for two cases: when the <strong>referenced table is partitioned</strong> (the probe would need to be routed to the right partition), and for <strong>temporal foreign keys</strong> (range containment needs more than a single lookup). Later fixes added more fallbacks: the referenced index has to be a btree, and its collation must match the column's.</p>
<p>If the answer is yes, <code>ri_FastPathCheck</code> does what the query used to do, by hand:</p>
<ol>
<li><p>Opens the parent table and its unique index.</p>
</li>
<li><p>Takes a snapshot. In the shipped version this happens <em>after</em> the table lock is granted. Taking it earlier opened a window where a parent row committed while you waited for the lock would be invisible, and the check would fail for a key that exists (<a href="https://github.com/postgres/postgres/commit/1390182683">commit 1390182683</a>).</p>
</li>
<li><p>Builds scan keys from the new row's values and probes the btree with <code>index_getnext_slot</code>.</p>
</li>
<li><p>Locks the row it found with <code>table_tuple_lock(..., LockTupleKeyShare, ...)</code> in <code>ri_LockPKTuple</code>. That's the same lock <code>FOR KEY SHARE</code> takes.</p>
</li>
<li><p>If the lock had to follow an update chain to reach the newest row version, <code>recheck_matched_pk_tuple</code> checks that the key still matches. That is the fast path's version of the recheck the executor does for <code>SELECT ... FOR UPDATE</code>.</p>
</li>
</ol>
<p>It also switches to the parent table owner's user id for the lookup, the same way the SPI path does, so permissions and row-level security behave the same.</p>
<h2>The benchmark</h2>
<p>A parent table with 1 million integer primary keys, and a child table with a foreign key to it. Each run inserts 1 million child rows in one <code>INSERT ... SELECT</code>, after a <code>TRUNCATE</code> and a <code>CHECKPOINT</code>. The same insert into a table with no foreign key is the baseline. Six runs per version; the charts show the median.</p>
<img src="https://cdn-images-1.medium.com/v2/resize:fit:1600/1*dM1cdgI2iGMzJujUfPwQSQ.png" alt="Median time to insert 1 million rows. Scattered keys: 18.6 2499 ms, 19 1509 ms, master 1473 ms, no foreign key 299 ms. Sequential keys: 18.6 1965 ms, 19 1021 ms, master 701 ms." style="display:block;margin:0 auto" />

<p><em>Median of 6 runs, 1 million rows per insert. Lower is better.</em></p>
<p>With scattered keys, the foreign key check itself (the time above the no-foreign-key baseline) drops from <strong>2.17 s in 18.6 to 1.21 s in 19</strong>. The whole insert is 1.66x faster.</p>
<h3>The control: turn the fast path off</h3>
<p>To check the gain really comes from the fast path, I ran the same test with a partitioned parent table. The fast path refuses partitioned parents, so 19 falls back to SPI:</p>
<img src="https://cdn-images-1.medium.com/v2/resize:fit:1600/1*UCB1CiPGEGwit3NOdK-4JA.png" alt="Partitioned referenced table: 18.6 7808 ms, 19 7395 ms." style="display:block;margin:0 auto" />

<p><em>Partitioned parent, scattered keys. The gap nearly disappears (about 5%).</em></p>
<p>Two things show up here. The gain disappears when the fast path is off, so the fast path is what made the difference. And a foreign key to a partitioned table is expensive either way: about 3.5x slower than the SPI path on a plain table, because every check has to find the right partition first.</p>
<h2>The batching that was pulled</h2>
<p>Three days after the fast path, a <a href="https://github.com/postgres/postgres/commit/b7b27eb41a5cc0b45a1a9ce5c1cde5883d7bc358">second commit</a> went further. Instead of probing the index once per row, it buffered up to 64 rows and checked them together. For single-column keys it passed all 64 values as one array scan key (<code>SK_SEARCHARRAY</code>), so the btree sorts them and walks its leaf pages once instead of descending from the root 64 times. The commit reported about 2.9x in total over the old code.</p>
<p>That batching is in the 19 betas but <strong>not in 19.0</strong>. On September 10 it was <a href="https://github.com/postgres/postgres/commit/25649d6e791c2d803fff0ce8b69ed968028a5edc">removed from REL_19_STABLE</a>. Buffered checks have to survive across trigger calls and be flushed at the right moment, and that state has to be right under nested triggers, subtransactions, deferred constraints and <code>SET CONSTRAINTS</code>. Miss one case and a buffered check never runs: the transaction commits a real foreign key violation without an error. The commit message says it plainly: with 19 close to release, there wasn't time to be sure every case was covered. Batching stays on master for a later release. That's a hard call to make that close to release, and it's the right one.</p>
<p>My first measurement disagreed with the 2.9x. With scattered keys, master (batching included) is only about 2% faster than 19. The array scan only saves work when the 64 values land on the same few leaf pages. My keys were deliberately spread over the whole index, so every value still needed its own leaf.</p>
<p>So I reran with the keys in order (<code>parent_id = 1, 2, 3, ...</code>), which is what you get when child rows are loaded in parent order. Now batching does what the commit said: <strong>701 ms against 1,021 ms</strong> for 19, and 2.8x against 18.6, close to the reported 2.9x.</p>
<table>
<thead>
<tr>
<th>Keys</th>
<th>18.6</th>
<th>19</th>
<th>master (batching)</th>
</tr>
</thead>
<tbody><tr>
<td>Scattered</td>
<td>2,499 ms</td>
<td>1,509 ms</td>
<td>1,473 ms</td>
</tr>
<tr>
<td>Sequential</td>
<td>1,965 ms</td>
<td>1,021 ms</td>
<td>701 ms</td>
</tr>
<tr>
<td>Scattered, partitioned parent</td>
<td>7,808 ms</td>
<td>7,395 ms</td>
<td>not run</td>
</tr>
</tbody></table>
<h2>What this means if you run Postgres</h2>
<ul>
<li><p><strong>Bulk loads into tables with foreign keys get faster in 19 for free</strong>, as long as the referenced table isn't partitioned and the key has a btree index (primary keys and unique constraints do).</p>
</li>
<li><p><strong>Foreign keys to partitioned tables don't get this.</strong> If you bulk-load into a table that references a partitioned table, the check still costs about 7 µs per row.</p>
</li>
<li><p><strong>Load order matters more after batching lands.</strong> Once batching ships in a later release, loading child rows sorted by parent key could make the check nearly twice as cheap again.</p>
</li>
<li><p><strong>Only the insert-side check changed.</strong> <code>ON DELETE CASCADE</code>, <code>SET NULL</code> and the other action triggers still go through SPI, as the original commit explains.</p>
</li>
</ul>
<h2>How I measured</h2>
<ul>
<li><p><strong>Machine:</strong> Apple M-series laptop. Each server ran in Docker (a Linux arm64 VM) with 4 CPUs and 4 GB of memory, one server at a time.</p>
</li>
<li><p><strong>Settings:</strong> <code>shared_buffers = 1GB</code>, <code>max_wal_size = 8GB</code> and <code>checkpoint_timeout = 30min</code> on every version. Everything else was left at the defaults.</p>
</li>
<li><p><strong>Versions:</strong></p>
<ul>
<li><p>18.6 and 19beta3 are the official Docker images.</p>
</li>
<li><p>"19" is REL_19_STABLE at <a href="https://github.com/postgres/postgres/commit/8fbbb338bed77c6fa8dc6fa309cbcca93c1c2d72">8fbbb338be</a>, which already has batching removed, built from source with <code>-O2</code>.</p>
</li>
<li><p>"master" is master at <a href="https://github.com/postgres/postgres/commit/4545cee">4545cee</a>, built the same way.</p>
</li>
<li><p>The 19beta3 image still includes batching and measured 1,529 ms on scattered keys, close to both 19 and master.</p>
</li>
</ul>
</li>
<li><p><strong>Caveats:</strong></p>
<ul>
<li><p>The source builds use a different compiler setup from the packaged images. That probably explains part of the small difference in the no-foreign-key baseline (332 ms against 299 ms).</p>
</li>
<li><p>These are single-session numbers. Concurrent inserts and lock contention are a separate test.</p>
</li>
</ul>
</li>
</ul>
<p>All scripts, Dockerfiles and raw results are in the <a href="https://getnobody.org/labs/pg19-fk-fastpath/">lab folder</a>. Run <code>./run.sh &lt;label&gt; &lt;image&gt;</code> and you get the same table.</p>
<p>Credit goes to Amit Langote, who wrote the fast path, the batching and roughly thirty follow-up commits to this code this year, and to everyone on pgsql-hackers who tested the betas hard enough to find the problems before release.</p>
<p><em>Originally published at</em> <a href="https://getnobody.org/blog/postgres-19-foreign-key-fast-path/"><em>getnobody.org</em></a><em>. Benchmark scripts and raw results:</em> <a href="https://getnobody.org/labs/pg19-fk-fastpath/"><em>getnobody.org/labs/pg19-fk-fastpath</em></a><em>.</em></p>
]]></content:encoded></item></channel></rss>