Bounds durable-agent heartbeat worktree acquisition to a fixed retry count instead of requeuing to todo indefinitely across heartbeat cycles. - Add MAX_HEARTBEAT_WORKTREE_ACQUISITION_RETRIES (3) in agent-heartbeat.ts, reusing Task.recoveryRetryCount as a cross-heartbeat counter (no schema migration) - On cap exhaustion, terminally mark the task status:"failed" with an explanatory error, log the entry, and reopen to todo with preserveStatus so the failed status isn't wiped by reopen-to-todo semantics - Add onTaskAcquisitionExhausted callback wired in in-process-runtime.ts to CentralCore.recordTaskCompletion(taskId, false) so exhausted acquisitions count toward totalTasksFailed - Add regression tests in agent-heartbeat-worktree.test.ts and in-process-runtime.test.ts covering the retry cap and completion recording - Add changeset (patch) and a docs/solutions/logic-errors writeup documenting the investigation and other worktree-collision sub-gaps found not to reproduce on HEAD Files changed: .changeset/fn-7721-worktree-heartbeat-retry-cap.md | 7 ++ docs/solutions/logic-errors/heartbeat-worktree-acquisition-unbounded-requeue.md | 84 ++++++++++++++++++++++ packages/engine/src/__tests__/agent-heartbeat-worktree.test.ts | 58 +++++++++++++++ packages/engine/src/__tests__/in-process-runtime.test.ts | 11 +++ packages/engine/src/agent-heartbeat.ts | 72 ++++++++++++++++++- packages/engine/src/runtimes/in-process-runtime.ts | 12 ++++ 6 files changed, 242 insertions(+), 2 deletions(-) Fusion-Task-Id: FN-7721 Fusion-Task-Lineage: caad671c-f360-4c1c-8aaa-5b48fca5a55b Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
4.0 KiB
title, date, category, module, problem_type, component, symptoms, root_cause, resolution_type, severity, related_components, tags
| title | date | category | module | problem_type | component | symptoms | root_cause | resolution_type | severity | related_components | tags | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Durable-agent heartbeat worktree acquisition retried unboundedly and failures went uncounted | 2026-07-09 | docs/solutions/logic-errors | engine agent heartbeat + worktree acquisition | logic_error | engine |
|
invariant_gap | code_fix | medium |
|
|
Durable-agent heartbeat worktree acquisition retried unboundedly and failures went uncounted
Problem
Executor.createWorktree (the main task-execution path) has always bounded its
worktree-creation retries via MAX_WORKTREE_RETRIES = 3 with exponential
backoff, and terminal failures flow through Executor's onError callback into
InProcessRuntime.recordTaskCompletion, which increments CentralCore's
totalTasksFailed.
HeartbeatMonitor.executeHeartbeat has a separate task-worktree-acquisition
call path used when a durable custom agent's heartbeat picks up its assigned
task (acquireTaskWorktree, distinct from Executor.createWorktree). Before
this fix, that path had no retry cap at all: on any acquisition failure it
unconditionally moved the task back to todo (preserveProgress: true) and
completed the heartbeat run successfully. Each subsequent heartbeat interval
was an independent, uncounted retry of the same acquisition — a persistently
failing collision (e.g. a branch genuinely owned by a live foreign task with
sibling-branch-rename disabled) could requeue to todo indefinitely across
many hours, and because the task never reached a terminal status: "failed"
state, the failure was never recorded via CentralCore.recordTaskCompletion.
Root Cause
Two independent retry/bounding mechanisms exist for worktree creation
(Executor.createWorktree's in-call loop, and the heartbeat's per-cycle call),
but only the executor path was ever wired to a shared retry-cap counter and to
CentralCore.recordTaskCompletion. The heartbeat path bypassed both.
Fix
agent-heartbeat.ts: addedMAX_HEARTBEAT_WORKTREE_ACQUISITION_RETRIES = 3. ReusesTask.recoveryRetryCount(no schema migration) as a cross-heartbeat counter. Below the cap, bump the counter and requeue as before. At/above the cap, mark the task terminallystatus: "failed"(the same convention the executor uses — the task stays visible intodoforfn_task_retry), log a clear error citing the branch and attempt count, and invoke a new optionalonTaskAcquisitionExhausted(taskId, detail)callback.runtimes/in-process-runtime.ts: wiresonTaskAcquisitionExhaustedtothis.recordTaskCompletion(taskId, false), so the failure is counted the same wayExecutor'sonErrorcounts one.
Prevention
When adding a new retry loop around a resource-acquisition call that already
has an established bounded-retry convention elsewhere in the codebase (here:
Executor.createWorktree's MAX_WORKTREE_RETRIES + NonRetryableWorktreeError
onError→recordTaskCompletion), verify the new call site reuses or mirrors that convention instead of independently reimplementing a "just requeue on failure" fallback with no cap and no failure-counting hook.
Related
docs/solutions/logic-errors/repo-root-task-worktree-requeue-loop.md— a different (already-fixed) unbounded-requeue shape in the executor's own resume path.- Regression tests:
packages/engine/src/__tests__/agent-heartbeat-worktree.test.ts,packages/engine/src/__tests__/in-process-runtime.test.ts.