Files
fusion/packages
gsxdsm dc132071d5 fix(#2411): let embedded PostgreSQL crash recovery finish instead of racing it
beta.4 follow-up from the issue thread: on an interrupted (not-cleanly-shut-down)
cluster, the elevated Windows launcher declared readiness on a bare TCP accept
while crash recovery still rejected every connection with 57P03, so
ensureDatabase failed and the start cleanup fast-shutdown the recovering
postmaster ~0.2s after launch; the retry then joined the instance it had just
told to stop, got ECONNREFUSED, and parked the dashboard in a dead shell. The
30s ".pgrunner sharing violation" stall was recovery's SyncDataDirectory fsync
walk hitting Fusion's own pgctl log inside the data dir.

- Move the pgctl runner dir to a sibling .pgrunner-<dataDirName> outside the
  data dir (and sweep the legacy in-dataDir .pgrunner), so recovery's fsync
  walk can never contend with the postmaster's inherited log handle.
- Ignore 57P03 recovery rejections in the elevated readiness fatal scan.
- Owned starts wait for the cluster to genuinely accept connections (retrying
  57P03/socket errors, bounded by the start timeout) before ensureDatabase —
  never stop a postmaster that is still in recovery.
- Join-path database verify retries the 57P03 recovery signal for up to 15s;
  socket errors keep the instant optimistic-join contract for stale pids.
- startup-factory's joined-instance-unreachable retry backs off across ~15s
  instead of a single 500ms attempt.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-23 09:13:21 -07:00
..
2026-07-23 00:16:34 -07:00
2026-07-23 00:16:34 -07:00
2026-07-23 00:16:34 -07:00
2026-07-23 00:16:34 -07:00
2026-07-23 00:16:34 -07:00
2026-07-23 00:16:34 -07:00
2026-07-23 00:16:34 -07:00
2026-07-23 00:16:34 -07:00