perf(backfill): fast-lane Phase-1 via lifo + run scan despite deep backlog #133

Merged
root merged 1 commits from fix/backfill-lifo-fast-lane into main 2026-06-12 11:39:41 +03:00
Owner

The hourly catalog-backfill scan had gone effectively dead on prod: its
job sat behind a ~430k-deep wait list and, even when it ran, self-skipped
because the queue backlog (431k) was far over the maxBacklog ceiling
(1000). Net result: newly-decoded / zero-parts vehicles were never
onboarded — they starved behind the deep-drill backlog (ic stuck at 28
for days, 0 backfill-scan jobs ever completed).

Root constraint (verified against bullmq 5.68 lua): moveToActive drains
the wait list (RPOPLPUSH from the tail) BEFORE the prioritized ZSET,
so a priority job is starved behind an already-deep wait queue — the
opposite of what's wanted. The lever that works is lifo: it RPUSHes to
the tail, where the very next RPOPLPUSH picks it, ahead of the FIFO
backlog. addJobFromScheduler honours lifo too, so the scan job itself can
jump the queue.

Changes:

  • Thread a fast flag through the init→children→parts chain; fast jobs
    are enqueued with lifo:true so the whole chain jumps the backlog.
  • Backfill scan: Phase-1 (zero-parts vehicles) now runs every wave in the
    fast lane even when the backlog is over the ceiling; only Phase-2 (the
    rolling rescan that piles on) is suspended while the queue is deep.
  • Register the hourly scan job with lifo:true so it fires on the next
    tick instead of being buried for days.

Tested: new prefetch-worker.service.spec (lifo wiring + Phase-1/Phase-2
gating), full api suite green (248 passed), typecheck + biome clean.

Co-Authored-By: Claude Fable 5 noreply@anthropic.com

The hourly catalog-backfill scan had gone effectively dead on prod: its job sat behind a ~430k-deep wait list and, even when it ran, self-skipped because the queue backlog (431k) was far over the maxBacklog ceiling (1000). Net result: newly-decoded / zero-parts vehicles were never onboarded — they starved behind the deep-drill backlog (ic stuck at 28 for days, 0 backfill-scan jobs ever completed). Root constraint (verified against bullmq 5.68 lua): moveToActive drains the `wait` list (RPOPLPUSH from the tail) BEFORE the `prioritized` ZSET, so a `priority` job is starved behind an already-deep wait queue — the opposite of what's wanted. The lever that works is `lifo`: it RPUSHes to the tail, where the very next RPOPLPUSH picks it, ahead of the FIFO backlog. addJobFromScheduler honours lifo too, so the scan job itself can jump the queue. Changes: - Thread a `fast` flag through the init→children→parts chain; fast jobs are enqueued with `lifo:true` so the whole chain jumps the backlog. - Backfill scan: Phase-1 (zero-parts vehicles) now runs every wave in the fast lane even when the backlog is over the ceiling; only Phase-2 (the rolling rescan that piles on) is suspended while the queue is deep. - Register the hourly scan job with `lifo:true` so it fires on the next tick instead of being buried for days. Tested: new prefetch-worker.service.spec (lifo wiring + Phase-1/Phase-2 gating), full api suite green (248 passed), typecheck + biome clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
root added 1 commit 2026-06-12 11:38:47 +03:00
perf(backfill): fast-lane Phase-1 via lifo + run scan despite deep backlog
Some checks failed
QA Gate (P0/P1) / Test affected app (pull_request) Has been cancelled
0de6bd3faa
The hourly catalog-backfill scan had gone effectively dead on prod: its
job sat behind a ~430k-deep wait list and, even when it ran, self-skipped
because the queue backlog (431k) was far over the maxBacklog ceiling
(1000). Net result: newly-decoded / zero-parts vehicles were never
onboarded — they starved behind the deep-drill backlog (ic stuck at 28
for days, 0 backfill-scan jobs ever completed).

Root constraint (verified against bullmq 5.68 lua): moveToActive drains
the `wait` list (RPOPLPUSH from the tail) BEFORE the `prioritized` ZSET,
so a `priority` job is starved behind an already-deep wait queue — the
opposite of what's wanted. The lever that works is `lifo`: it RPUSHes to
the tail, where the very next RPOPLPUSH picks it, ahead of the FIFO
backlog. addJobFromScheduler honours lifo too, so the scan job itself can
jump the queue.

Changes:
- Thread a `fast` flag through the init→children→parts chain; fast jobs
  are enqueued with `lifo:true` so the whole chain jumps the backlog.
- Backfill scan: Phase-1 (zero-parts vehicles) now runs every wave in the
  fast lane even when the backlog is over the ceiling; only Phase-2 (the
  rolling rescan that piles on) is suspended while the queue is deep.
- Register the hourly scan job with `lifo:true` so it fires on the next
  tick instead of being buried for days.

Tested: new prefetch-worker.service.spec (lifo wiring + Phase-1/Phase-2
gating), full api suite green (248 passed), typecheck + biome clean.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
root merged commit cbdff3311f into main 2026-06-12 11:39:41 +03:00
root deleted branch fix/backfill-lifo-fast-lane 2026-06-12 11:39:52 +03:00
Sign in to join this conversation.
No Reviewers
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: root/sase.tr#133