Investigation of reported DB read contention found two cross-process
contention sources in the SQLite layer (single synchronous node:sqlite
connection per process, WAL mode):
- Unbounded WAL on central-db and archive-db. Neither set
journal_size_limit, so their WAL never truncated back down after a
checkpoint and every reader paid an ever-growing WAL-index scan. Add
journal_size_limit=4MB (matching db.ts) plus explicit
synchronous=FULL/wal_autocheckpoint=1000 for intent. central-db is the
most cross-process-shared DB; archive-db had the same latent gap.
- vacuum() held the EXCLUSIVE lock past its own runtime. Resetting
locking_mode to NORMAL does not drop the WAL exclusive lock until the
connection next touches the DB, so other processes stayed locked out of
reads (SQLITE_BUSY) until some unrelated query ran. A plain read does
NOT release it in WAL mode (verified); a PASSIVE checkpoint does. Run
one in the finally, guard the locking_mode reset so it can't mask the
original error or skip the release, and log swallowed failures.
Tests: assert the new PRAGMAs on central-db and archive-db, and that a
second connection can read immediately after vacuum() returns.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>