Files
fusion/packages/dashboard/app/hooks/resyncRetry.ts
gsxdsm f26cbedf4f fix(dashboard): close the code-review findings on the mobile tab-discard work
An 11-reviewer pass over f157bf7460..f5163d8351 found defects in the mobile
tab-discard change set itself. This fixes them.

Silent data loss (the recurring defect class):
- AgentDetailView reconnect refetched limit:100 and replaced wholesale, so 380
  displayed lines vanished with no "Load older" and no indicator; it now
  reconciles through the shared logStreamReconcile helper.
- useActivityLog.loadMore past the cap discarded the page it had just fetched
  while advancing the cursor and leaving hasMore true, so the feed silently
  stopped paginating behind a live-looking button.
- useAgentLogs: loadMore and resyncFromServer had no mutual exclusion, a
  no-overlap resync discarded explicitly paged-back history, a resync outliving
  the reconnect delay left an unmarked gap, and the live-tail trim could evict
  the gap marker itself.
- useLiveTranscript's resync overwrote live entries that raced the refetch.

The premise itself was not fully delivered:
- useProjects, useNodes, and useMeshState never called clearInterval, so they
  polled the whole time the tab was hidden. useProjects is mounted for the
  entire session, so the page never went idle -- the primary mechanism this
  work depends on. All three now use the shared visibility gate.
- sse-bus fired onReconnect twice per reconnect cycle and fanned out ~28
  subscribers in one tick, against a ~6-connection-per-origin cap on a waking
  radio. The successful open is now the single authority, and the fan-out uses
  the same exported stagger primitive as the polling path rather than a second
  copy of the slot formula.
- A channel first subscribed during the hidden window opened a live EventSource
  and keepalive; suspension is now a module-level condition openChannel
  consults, and a channel opened inside the grace window re-arms it.

Credentials and correctness:
- The service worker persisted every GET /api/* to durable Cache Storage,
  including /api/settings with daemonToken, githubAuthToken, gitlabAuthToken
  and ntfyAccessToken in plaintext, with no exclusion and no purge path --
  "Clear all cached data" only walked localStorage. Now gated, bounded, and
  genuinely purgeable.
- useTasks cleared its own snapshot when the mount revalidation failed on a
  waking radio, so the board blanked and the next restore was empty too.
  Suspension-class failures no longer destroy the cache.
- A single-row SSE update reset lastFetchTimeMs to now while an hours-old
  hydrated snapshot was on screen, re-marking every in-progress card stuck.
- ListView's "Select all visible tasks" acted on the full filtered set while
  only 50 rows rendered, so a bulk delete reached rows the operator could not
  see. Column's search window reset keyed on a boolean, so refining a query
  kept the expanded window.

Tests that could not fail:
- App.test.tsx mocked TerminalModal as isOpen ? <div/> : null, making the
  unmount-on-close invariant unobservable; MockEventSource kept its listeners
  after close(), so cases passed with their onReconnect handlers deleted.
- The SSE resync ratchet scanned only hooks/, exempting ~13 component call
  sites -- the exact regression it exists to prevent.
- MissionControlPanel's bespoke poll and the xterm scrollback constants and
  WebGL disposal had no coverage at all.

Verified: tsc -p tsconfig.app.json clean, pnpm lint clean, pnpm
check:changesets clean, 877 tests passing across 36 scoped files.
Known unrelated red: MailboxView.test.tsx's FN-8407 CSS guard fails at HEAD
too -- this diff adds no @media rule and no .mailbox-view--mobile selector,
the only two things that assertion inspects. Left alone deliberately.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 11:17:52 -07:00

117 lines
4.3 KiB
TypeScript

/*
FNXC:SseResync 2026-07-26-18:05:
ONE bounded retry ladder for every "the SSE channel just reopened, refetch what we missed" handler.
Why it exists: those handlers were each written as `void refetch().catch(() => {})` with a comment
saying the next reconnect retries. That claim is false in the case that matters — after a mobile
hidden-tab suspend the stream reopens ONCE, and if that single refetch fails (the same flaky network
that caused the suspend) the surface stays silently wrong until the operator reloads. For the
approval banner that means a pending decision is never rendered and the requesting agent blocks
indefinitely.
Bounded, not unbounded: the resume edge is already congested (every suspended channel refetches at
the same instant), so a failing surface gets a short ladder and then STOPS and reports degradation
via `onExhausted` — it does not hammer the server.
Shared rather than copied: three near-identical hand-rolled fixes in one change set is the documented
recurring failure mode here (AGENTS.md "Reuse Components, Design Tokens, and Systems (No Drift)"), so
every resync handler imports this runner instead of growing its own variant.
*/
/** Delays before retry attempts 2 and 3. Two retries ≈ 10s of coverage for a transient failure. */
export const DEFAULT_RESYNC_RETRY_DELAYS_MS: readonly number[] = [2_000, 8_000];
export interface ResyncRetryOptions {
/** One resync attempt. MUST reject on failure — a promise that swallows its own error is not retryable. */
run: () => Promise<void>;
/** Delay before each retry. Length bounds the number of retries (default: two). */
delaysMs?: readonly number[];
/**
* Every attempt in the ladder failed. The surface's data is now KNOWN to be possibly incomplete;
* callers use this to raise an operator-visible signal rather than render a confident wrong view.
*/
onExhausted?: (error: unknown) => void;
/** An attempt succeeded after a previous ladder had been exhausted. Callers clear their signal. */
onRecovered?: () => void;
}
export interface ResyncRetryRunner {
/** Run the resync now, cancelling any pending retry (a fresh reconnect supersedes a scheduled one). */
trigger: () => void;
/** Cancel pending retries and suppress all further callbacks. Call from effect cleanup. */
dispose: () => void;
}
export function createResyncRetryRunner(options: ResyncRetryOptions): ResyncRetryRunner {
const delays = options.delaysMs ?? DEFAULT_RESYNC_RETRY_DELAYS_MS;
let timer: ReturnType<typeof setTimeout> | null = null;
let inFlight = false;
let disposed = false;
// True once a full ladder failed; drives the single onRecovered edge.
let exhausted = false;
const clearTimer = () => {
if (timer !== null) {
clearTimeout(timer);
timer = null;
}
};
const onFailure = (retryIndex: number, error: unknown) => {
inFlight = false;
if (disposed) return;
if (retryIndex < delays.length) {
clearTimer();
timer = setTimeout(() => {
timer = null;
attempt(retryIndex + 1);
}, delays[retryIndex]);
return;
}
exhausted = true;
options.onExhausted?.(error);
};
const onSuccess = () => {
inFlight = false;
if (disposed) return;
if (exhausted) {
exhausted = false;
options.onRecovered?.();
}
};
function attempt(retryIndex: number): void {
// Overlapping attempts are pointless: they converge on the same server state and double the
// load on the resume edge this ladder is trying not to congest.
if (disposed || inFlight) return;
inFlight = true;
/*
`run()` is invoked SYNCHRONOUSLY, never deferred through a microtask. Callers set their
"resync in flight — park live events" flag in the first synchronous statement of run(), and an
event delivered during an intervening microtask would slip past that flag and then be overwritten
by the page the resync applies: exactly the silent data loss these resyncs exist to prevent.
*/
let started: Promise<void>;
try {
started = options.run();
} catch (error: unknown) {
onFailure(retryIndex, error);
return;
}
void started.then(onSuccess, (error: unknown) => onFailure(retryIndex, error));
}
return {
trigger: () => {
if (disposed) return;
clearTimer();
attempt(0);
},
dispose: () => {
disposed = true;
clearTimer();
},
};
}