Commit Graph

1714 Commits

Author SHA1 Message Date
gsxdsm
0d0d818b77 feat(FN-4891): merge fusion/fn-4891 2026-05-17 08:24:49 -07:00
Fusion (runfusion.ai)
da6598a4c0 fix(FN-4759): complete Step 5-6 — pass gates and add changeset
Fusion-Task-Id: FN-4759
Fusion-Task-Lineage: af79d2a6-fd7b-40e0-b3c9-6c74569d9aa0
2026-05-17 08:22:26 -07:00
Fusion (runfusion.ai)
96a19309ae fix(FN-4884): finalize docs and keep workspace gates green
Fusion-Task-Id: FN-4884
Fusion-Task-Lineage: b906ade4-a9f1-4f2e-b684-4ebf49c3e7e6
2026-05-17 07:26:21 -07:00
Fusion (runfusion.ai)
030a27b40d feat(FN-4884): complete Step 2-3 fast-path and coverage
Fusion-Task-Id: FN-4884
Fusion-Task-Lineage: b906ade4-a9f1-4f2e-b684-4ebf49c3e7e6
2026-05-17 07:26:21 -07:00
gsxdsm
83ed893382 feat(FN-4879): merge fusion/fn-4879 2026-05-17 06:23:00 -07:00
Fusion (runfusion.ai)
75ac66a8a0 test(FN-4866): complete Step 2 — add named duplicate-search regression case
Fusion-Task-Id: FN-4866
Fusion-Task-Lineage: 2188be29-febc-4610-92f6-d7c1a911359e
2026-05-17 02:46:46 -07:00
gsxdsm
ee4652a492 feat(FN-4861): merge fusion/fn-4861 2026-05-17 02:08:01 -07:00
gsxdsm
404e202d7e feat(FN-4809): merge fusion/fn-4809 2026-05-17 00:53:02 -07:00
gsxdsm
3ca3b5e471 feat(FN-4851): merge fusion/fn-4851 2026-05-17 00:38:07 -07:00
Fusion
aa6d1c9851 fix(FN-4847): discard foreign branch and recreate on branch-conflict-unrecoverable
Production failure shape:
  Auto-recovery failed: branch conflict unrecoverable \u2014
  Branch fusion/fn-4847 is already checked out at /.../deft-crane
  (tip a881ccc86660, 24 stranded commits since 0b28388876).
  Run branch recovery and explicitly choose whether to reclaim or
  discard prior work.

The 24 stranded commits are cross-task contamination residue from the
FN-4781/FN-4804/FN-4814 worktree-race era \u2014 they are NOT FN-4847's work.
Previously this paused the task with pausedReason='branch-conflict-
unrecoverable' and the task got stuck forever waiting for human
adjudication.

User intent (FN-4847): 'just create a new branch and keep going and
discard the old one'. Implementation:

1. auto-recovery.ts:actionForMode \u2014 in 'deterministic-only' mode (the
   default), branch-conflict-unrecoverable now returns 'retry' (was
   'pause'). This routes the failure to the handler instead of pausing.

2. auto-recovery-handlers/branch-worktree.ts \u2014 'live-foreign' inspection
   no longer emits irreducible-pause. Instead:
   - Check FN-4811 active-session registry. If the foreign worktree is
     bound to a live executor/merger session, do NOT force-remove it
     (would yank the live agent's filesystem). Just requeue and let
     downstream conflict-recovery handle it.
   - Otherwise: force-delete the foreign worktree (--force) + prune git
     worktree admin entries + force-delete the branch. Errors at each
     step are best-effort and logged.
   - Emit new audit event 'branch-worktree:foreign-branch-discarded'
     with stranded-commit count, live-ownership flag, success flags.
   - Requeue task to 'todo' with preserveProgress, clearing
     branch+baseCommitSha.

3. run-audit.ts \u2014 register new DatabaseMutationType.

4. executor-worktree.test.ts \u2014 update the 'records recovery context'
   test to assert the new retry+requeue contract (was asserting the old
   pause-with-status-failed contract).

Verification:
  - Targeted suite (4 files, 343 tests): pass.
  - pnpm --filter @fusion/engine build: clean.
  - pnpm lint: clean.

Fusion-Task-Id: FN-4847
2026-05-16 23:53:56 -07:00
Fusion (runfusion.ai)
01f0bf625a fix(FN-4826): complete Step 6 — stabilize audit emission and tests
Fusion-Task-Id: FN-4826
Fusion-Task-Lineage: b08a58c8-66ab-43d4-bdf9-69d2b56b7bcc
2026-05-16 23:45:17 -07:00
Fusion (runfusion.ai)
1553f1dd0d test(FN-4826): complete Step 4 — cover handoff telemetry paths
Fusion-Task-Id: FN-4826
Fusion-Task-Lineage: b08a58c8-66ab-43d4-bdf9-69d2b56b7bcc
2026-05-16 23:45:17 -07:00
Fusion (runfusion.ai)
3453e77396 feat(FN-4826): complete Step 3 — emit mesh lease handoff telemetry
Fusion-Task-Id: FN-4826
Fusion-Task-Lineage: b08a58c8-66ab-43d4-bdf9-69d2b56b7bcc
2026-05-16 23:45:17 -07:00
Fusion (runfusion.ai)
56aca02250 feat(FN-4826): complete Step 2 — emit scheduler handoff events
Fusion-Task-Id: FN-4826
Fusion-Task-Lineage: b08a58c8-66ab-43d4-bdf9-69d2b56b7bcc
2026-05-16 23:45:17 -07:00
Fusion (runfusion.ai)
0705f42a38 feat(FN-4826): complete Step 1 — declare mutation types
Fusion-Task-Id: FN-4826
Fusion-Task-Lineage: b08a58c8-66ab-43d4-bdf9-69d2b56b7bcc
2026-05-16 23:45:17 -07:00
Fusion (runfusion.ai)
3605c67925 feat(FN-4791): add core secrets store for secure key-value storage
Adds a new secrets store module to `@fusion/core` (277 lines in `secrets-store.ts`) and exports it from the package index.

Fusion-Task-Id: FN-4791

Fusion-Task-Lineage: 7ff20b8a-37e1-46c6-8003-b542df9f98b1
2026-05-16 23:44:47 -07:00
Fusion (runfusion.ai)
e1667a789d test(FN-4815): complete Step 2 — add duplicate-search regression coverage
Fusion-Task-Id: FN-4815
Fusion-Task-Lineage: dc29e38d-5914-40c2-b9f7-3504606d91a3
2026-05-16 23:31:56 -07:00
Fusion (runfusion.ai)
e3aeec999c feat(FN-4815): complete Step 1 — add duplicate scenario fixture
Fusion-Task-Id: FN-4815
Fusion-Task-Lineage: dc29e38d-5914-40c2-b9f7-3504606d91a3
2026-05-16 23:31:56 -07:00
Fusion (runfusion.ai)
8c3c0206f1 feat(FN-4836): consolidate settings round-trip/default tests with it.each m
Refactors the settings store tests in `packages/core` by consolidating round-trip and default value coverage into structured `it.each` matrices, trimming the test file by roughly 60 lines while maintaining equivalent coverage.

Fusion-Task-Id: FN-4836
2026-05-16 22:47:24 -07:00
Fusion (runfusion.ai)
5cd0e66654 feat(FN-4839): harden real-git test timeouts across engine suites
Hardens test timeouts across 12 reliability and integration test files (branch-conflict recovery, merger diff/overlap guards, self-healing, worktree hydration, workflow/file-scope interactions), covering both real-git and mock-based test lanes with consistent timeout adjustments to reduce flakiness.

Fusion-Task-Id: FN-4839
2026-05-16 22:36:57 -07:00
Fusion (runfusion.ai)
1aeb070083 test(FN-4824): reliability-interaction coverage for cross-node assignment wakes
Fusion-Task-Id: FN-4824
Fusion-Task-Lineage: 65f53a33-0939-4b6f-b62b-4a255ad14d9e
2026-05-16 21:53:18 -07:00
Fusion (runfusion.ai)
46a50c123f feat(FN-4824): forward task assignment events across remote runtime
Fusion-Task-Id: FN-4824
Fusion-Task-Lineage: 65f53a33-0939-4b6f-b62b-4a255ad14d9e
2026-05-16 21:53:18 -07:00
Fusion (runfusion.ai)
70728db098 fix(FN-4834): align reliability fixture task fields with task typing
Fusion-Task-Id: FN-4834
Fusion-Task-Lineage: 714c9c73-ca66-4c3f-980c-d03e6b50316a
2026-05-16 21:34:25 -07:00
Fusion (runfusion.ai)
594ec3e515 test(FN-4834): complete Step 2 — add reliability init stderr regression coverage
Fusion-Task-Id: FN-4834
Fusion-Task-Lineage: 714c9c73-ca66-4c3f-980c-d03e6b50316a
2026-05-16 21:34:25 -07:00
Fusion (runfusion.ai)
ed8c8752f0 feat(FN-4834): complete Step 1 — surface init diagnostics in task log
Fusion-Task-Id: FN-4834
Fusion-Task-Lineage: 714c9c73-ca66-4c3f-980c-d03e6b50316a
2026-05-16 21:34:25 -07:00
Fusion (runfusion.ai)
405d24f9ab fix(FN-4823): tighten lease recovery audit semantics
Fusion-Task-Id: FN-4823
Fusion-Task-Lineage: 0034c04f-df82-4a1b-9f25-b93643a2c157
2026-05-16 21:26:04 -07:00
Fusion (runfusion.ai)
a9dfb17b3d test(FN-4823): expand central-claim recovery coverage
Fusion-Task-Id: FN-4823
Fusion-Task-Lineage: 0034c04f-df82-4a1b-9f25-b93643a2c157
2026-05-16 21:26:04 -07:00
Fusion (runfusion.ai)
bbcc0a269c feat(FN-4823): wire central claim store into mesh lease manager
Fusion-Task-Id: FN-4823
Fusion-Task-Lineage: 0034c04f-df82-4a1b-9f25-b93643a2c157
2026-05-16 21:26:04 -07:00
Fusion (runfusion.ai)
6fa9aef630 feat(FN-4823): complete lease recovery central-claim reconciliation
Fusion-Task-Id: FN-4823
Fusion-Task-Lineage: 0034c04f-df82-4a1b-9f25-b93643a2c157
2026-05-16 21:26:04 -07:00
Fusion
5c36c0f15f fix(FN-4811): scope-leak guard always allows .changeset/ paths
The [scope-leak] reviewLevel=N enforcement=warn warning was firing on
many in-progress tasks for off-scope .changeset/FN-XXXX-*.md files (the
production signature on FN-4789, FN-4801, FN-4818 \u2014 their branches all
contained .changeset/FN-4811-*.md files from the in-progress fix stack).

By convention every task may add its own changeset entry under
.changeset/ per AGENTS.md 'Finalizing Changes' section, so .changeset/
files are now treated as always-allowed by the scope-leak guard
regardless of the task's declared file scope.

Cross-task changeset leakage is still caught by stronger downstream
guards (file-scope invariant at squash, post-merge audit) at much
higher signal-to-noise. This change only suppresses the noisy
per-execution warning that was flooding logs without adding any
defensive value.

Adds a new exported helper isAlwaysAllowedScopeLeakPath() so the
allowlist surface is easy to extend. Test coverage in
scope-leak-changeset-allowlist.test.ts.

Also (incorporated from interrupted merge state): loosens the
executing-task-lock.test.ts assertion that one losing-instance store
sees zero work-log entries rather than the brittle exact-count of
mockedCreateFnAgent invocations (the no-fn_task_done retry path can
fire on the winning instance, so the count varies).

Fusion-Task-Id: FN-4811
2026-05-16 21:15:10 -07:00
Fusion
b6df11a6a8 fix(FN-4811): process-wide executingTaskLock blocks parallel execute() across instances
Investigating FN-4814 + FN-4811 re-failures after commit 8bef30655 (which
added per-instance synchronous this.executing.add) revealed the per-instance
guard was insufficient. FN-4809 log at 02:48:17-18 UTC:

  02:48:17  [-]                Resuming execution after unpause
  02:48:17  [-]                Step 4 (Testing & Verification) -> pending
  02:48:17  [-]                Step 4 (Testing & Verification) -> pending
  02:48:17  [6097725-y2nb]     Executor detected stale merge state ...
  02:48:18  [6097816-9gde]     Executor detected stale merge state ...

Both runs y2nb and 9gde reached executor.ts:2661 (which is INSIDE execute(),
past the synchronous this.executing.add claim). The only viable explanation
is that there is more than one TaskExecutor instance in the same Node
process (engine restart race, multi-project hybrid runtime, or similar code
path). Each instance has its own executing Set, so the per-instance guard
doesn't help.

Fix: module-level singleton executingTaskLock in active-session-registry.ts,
shared across all TaskExecutor instances. execute() synchronously tryClaim()s
the lock; if false, bails. Every existing this.executing.delete() site also
calls executingTaskLock.release(). Per-instance this.executing kept because
many other call sites use it (this.executing.has at handler gates,
stuck-detector, resumeTaskForAgent, etc.).

Test setup (resetExecutorMocks in executor-test-helpers.ts) clears the lock
between tests so process-wide state doesn't leak (executor-pause and
executor-prompt tests would otherwise show 'expected 2 createFnAgent calls
but got 0' / 'expected not called but called 3 times' flakes).

Tests:
  - executing-task-lock.test.ts: 2 cases. Key case creates TWO TaskExecutor
    instances and races them on the same task ID, asserts only ONE actually
    runs. Verified FAILS on prior code (8bef30655) and PASSES on fix.

Verification:
  - Targeted suite (4 files, 170 tests): pass.
  - pnpm --filter @fusion/engine build: clean.
  - pnpm lint: clean.

Fusion-Task-Id: FN-4811
2026-05-16 21:06:51 -07:00
Fusion (runfusion.ai)
27dd927213 feat(FN-4830): complete Steps 4-7 stale lock recovery delivery
Fusion-Task-Id: FN-4830
Fusion-Task-Lineage: d9b8ad72-669f-488e-8e85-be2dce9b8341
2026-05-16 20:54:39 -07:00
Fusion (runfusion.ai)
161cb565e6 feat(FN-4830): complete Step 3 — executor stale-lock recovery
Fusion-Task-Id: FN-4830
Fusion-Task-Lineage: d9b8ad72-669f-488e-8e85-be2dce9b8341
2026-05-16 20:54:34 -07:00
Fusion (runfusion.ai)
753068482e feat(FN-4830): complete Step 2 — backend stale-lock recovery
Fusion-Task-Id: FN-4830
Fusion-Task-Lineage: d9b8ad72-669f-488e-8e85-be2dce9b8341
2026-05-16 20:54:33 -07:00
Fusion (runfusion.ai)
9e2a2e37d0 feat(FN-4830): complete Step 1 — add stale lock helper
Fusion-Task-Id: FN-4830
Fusion-Task-Lineage: d9b8ad72-669f-488e-8e85-be2dce9b8341
2026-05-16 20:54:33 -07:00
Fusion (runfusion.ai)
a11bd0719e feat(FN-4822): complete Step 8 — document central claim authority
Fusion-Task-Id: FN-4822
Fusion-Task-Lineage: 08cc29e8-114a-48dc-80de-8d7fd2ce0e69
2026-05-16 20:48:00 -07:00
Fusion (runfusion.ai)
199f317813 feat(FN-4822): complete Step 6 — enforce central-claim race coverage
Fusion-Task-Id: FN-4822
Fusion-Task-Lineage: 08cc29e8-114a-48dc-80de-8d7fd2ce0e69
2026-05-16 20:48:00 -07:00
Fusion (runfusion.ai)
ad4b236985 test(FN-4822): complete Step 5 — add cross-node claim mutex integration race test
Fusion-Task-Id: FN-4822
Fusion-Task-Lineage: 08cc29e8-114a-48dc-80de-8d7fd2ce0e69
2026-05-16 20:48:00 -07:00
Fusion (runfusion.ai)
bb9e93faba fix(FN-4818): align owning-node interaction assertions with landed behavior
Fusion-Task-Id: FN-4818
Fusion-Task-Lineage: 1619844a-3945-4681-af58-ca32e5533514
2026-05-16 20:27:40 -07:00
Fusion (runfusion.ai)
5c5b922d79 test(FN-4818): complete Step 2 — add owning-node handoff interaction coverage
Fusion-Task-Id: FN-4818
Fusion-Task-Lineage: 1619844a-3945-4681-af58-ca32e5533514
2026-05-16 20:27:40 -07:00
Fusion (runfusion.ai)
003b343b4e test(FN-4818): complete Step 1 — add claim mutex interaction coverage
Fusion-Task-Id: FN-4818
Fusion-Task-Lineage: 1619844a-3945-4681-af58-ca32e5533514
2026-05-16 20:27:40 -07:00
Fusion (runfusion.ai)
2617ae659f test(FN-4818): complete Step 1 — add multi-node claim mutex interaction coverage
Fusion-Task-Id: FN-4818
Fusion-Task-Lineage: 1619844a-3945-4681-af58-ca32e5533514
2026-05-16 20:27:39 -07:00
Fusion (runfusion.ai)
1c7d49c5eb test(FN-4827): recover FN-4774 triage duplicate-detection regression block (orphan fusion/fn-4817; supersedes FN-4815) 2026-05-16 20:02:35 -07:00
Fusion (runfusion.ai)
14cf263bc6 feat(FN-4825): complete Step 3 — add scheduler unreachable-owner audits
Fusion-Task-Id: FN-4825
Fusion-Task-Lineage: dc4e633e-fe46-47a8-9ecc-f032073caee9
2026-05-16 19:58:44 -07:00
Fusion (runfusion.ai)
ee121a9f79 feat(FN-4825): complete Step 2 — emit mesh lease unreachable-owner audits
Fusion-Task-Id: FN-4825
Fusion-Task-Lineage: dc4e633e-fe46-47a8-9ecc-f032073caee9
2026-05-16 19:58:44 -07:00
Fusion (runfusion.ai)
2db5c10f07 feat(FN-4825): complete Step 1 — extend run-audit taxonomy
Fusion-Task-Id: FN-4825
Fusion-Task-Lineage: dc4e633e-fe46-47a8-9ecc-f032073caee9
2026-05-16 19:58:44 -07:00
Fusion (runfusion.ai)
ea6501bed7 test(FN-4810): complete Step 2 — cover truncation behavior
Fusion-Task-Id: FN-4810
Fusion-Task-Lineage: b7080f83-96b0-4106-9742-e14b20afeab1
2026-05-16 18:57:58 -07:00
Fusion (runfusion.ai)
af4607ed54 feat(FN-4810): complete Step 1 — truncate scope-leak log rendering
Fusion-Task-Id: FN-4810
Fusion-Task-Lineage: b7080f83-96b0-4106-9742-e14b20afeab1
2026-05-16 18:57:58 -07:00
Fusion
8bef30655d fix(FN-4811): close concurrent-execute race that produced parallel runs
TaskExecutor.execute() had a classic JS async race window. Original:

  async execute(task) {
    if (this.executing.has(task.id)) return;                       // check
    const assignedAgentId = task.assignedAgentId;
    if (assignedAgentId && await this.shouldDeferForHeartbeat(...)) // AWAIT yields
      return;
    this.executing.add(task.id);                                   // add (too late)
    ...
  }

Two concurrent execute(task) calls (scheduler dispatch + task:moved event
handler + restart-recovery) both:
  1. Pass the synchronous has() check (Set is empty).
  2. Enter the awaited shouldDeferForHeartbeat call (yields the event loop).
  3. Resume and both call this.executing.add(task.id).
  4. Both proceed to create the same worktree path.

Production failure shape (FN-4814 + FN-4811, observed within minutes):

  01:30:56  [runA-caoe]  Worktree created at /...worktrees/bright-mesa
  01:30:56  [runB-w23q]  Worktree created at /...worktrees/bright-mesa
  01:30:58              worktree liveness assertion failed: not_usable_task_worktree
  01:31:48              [thirdRun] also fires liveness assertion fail
  01:37:48              In-review stall surfaced [no-worktree-no-merge-confirmed]

This is the root cause of the entire FN-4781/FN-4804/FN-4814/FN-4811
cascade. Every other guard added today (FN-4811 active-session gate,
self-healing reclaim defer, validation-failed recovery, silent reclaim
recovery, integrity-warning dedup) was patching SYMPTOMS of the
duplicate-run race. With this fix, the symptoms stop appearing.

Fix: claim the slot synchronously immediately after the has() check,
release it on the heartbeat-defer early-return path. No await happens
between check and claim, so the race window is closed.

Test added under
packages/engine/src/__tests__/reliability-interactions/concurrent-execute-race.test.ts
verified to fail on the prior (a1b1f9aa0) executor.ts and pass on the
fixed version:

  - Two concurrent execute() calls produce the SAME number of
    createFnAgent invocations as one execute() call (no amplification).
  - A second sequential execute() after the first completes IS allowed
    (slot was released).

The task must have assignedAgentId set to exercise the race \u2014 without
it, the short-circuit `assignedAgentId && ...` evaluates the left side
to false synchronously, and no await happens.

Full engine suite: 5048+ tests pass. The 7 transient test-file failures
in the broad parallel run are pre-existing flaky real-git tests
(branch-conflicts-zero-unique, branch-conflicts-recovery,
merger-overlap-guard subprocess-guard contention) \u2014 all of them pass
when run alone or as a smaller group, none touch the executor.execute()
path.

Fusion-Task-Id: FN-4811
2026-05-16 18:55:03 -07:00
Fusion
a1b1f9aa05 fix(FN-4811): defer self-healing reclaim when worktree has an active session
The reclaimSelfOwnedBranchConflicts sweep was force-pausing actively-running
tasks. Production failure shape on FN-4819:

  1. Self-healing sweep runs every cycle and inspects branch conflicts.
  2. For FN-4819, inspection classified the conflict as 'tip-already-merged'
     (the task's branch tip was already on main).
  3. Sweep called removeWorktree({ reason: SelfHealingBranchConflict }).
  4. The FN-4811 active-session gate correctly refused: the worktree was
     bound to FN-4819/executor (a live agent session was using it).
  5. The thrown ActiveSessionWorktreeRemovalError was caught by the outer
     reclaim catch block.
  6. The catch escalated to AutoRecoveryDispatcher with class
     'branch-conflict-unrecoverable'.
  7. decision.action === 'pause' marked the task failed + paused +
     pausedReason='branch-conflict-unrecoverable' + moved to in-review.

Net effect: the FN-4811 gate (which is correct \u2014 you can't yank a live
worktree) became a regression source because the self-healing sweep
interpreted the refusal as fatal. Tasks that were actively making progress
got paused with a misleading 'branch conflict unrecoverable' error.

Fix: at the top of the per-task reclaim loop in
reclaimSelfOwnedBranchConflicts, check
activeSessionRegistry.isPathActive(task.worktree) and continue for any
task whose worktree is currently bound to a live session. The reclaim
will retry on the next sweep (sweeps run every cycle) once the session
has finished using the worktree. No data is lost, no decision is forced.

Test added under
packages/engine/src/__tests__/reliability-interactions/reclaim-defers-on-active-session.test.ts
covering:
  - The skip path: when activeSessionRegistry has a registration for
    task.worktree, the sweep MUST NOT call inspectBranchConflict,
    removeWorktree, or isUsableTaskWorktree. The task MUST stay in
    in-progress, not be marked failed/paused, not be moved to in-review.
  - Control: with no registration, the sweep DOES proceed and reaches
    inspectBranchConflict (preserving existing behavior).

Full engine suite: 314 files, 5061 tests pass, 1 skipped. Lint clean.
Build clean.

Fusion-Task-Id: FN-4811
2026-05-16 17:59:03 -07:00