Commit Graph

880 Commits

Author SHA1 Message Date
gsxdsm
cdf81a2244 fix(FN-8826): retry partial workflow progress after restart
Fusion-Task-Id: FN-8826
2026-08-07 16:55:00 -07:00
gsxdsm
ef8828f145 fix: unblock workflow execution stalled by silent role-routing deadlocks
Every task sat in progress with no session, no log, and no error after the
FN-8764 role-agent rollout. Two independent deadlocks, both invisible:

1. The in-process runtime built its AgentStore but never passed it into
   TaskExecutorOptions, so the executor's fail-closed role-routing gate refused
   every classified node (execute/step-execute/review/merge).
2. A resumed run keeps the continuation work item it woke on active until the
   interpreter returns, so the next node's principal-fence upsert violated
   idx_workflow_work_items_one_active_task_continuation — a different index than
   its ON CONFLICT target — and raised. The run re-suspended on every dispatch;
   only an operator bouncing the card to the hold column cleared it.

Both refusals were swallowed as recoverable "principal holds" that write no log,
audit row, or task error, which is why a fully deadlocked board looked idle.

- Wire agentStore into the executor; assert the shared instance at every runtime
  seam in the PG composition test.
- Supersede an active work item for a node the run has already left, then retry
  the fence write once; never touch a claim on the node currently executing.
- Record task:workflow-run-suspended and task:workflow-continuation-superseded;
  log principal holds, routing-unavailable faults, and fence-write errors.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 16:15:39 -07:00
gsxdsm
eaadd153b1 FN-8764: route workflow stages through durable role agents
Route workflow stages through task-scoped durable role agents.

- Persist normalized multi-role agents and workflow principal fences with migrations.
- Route planning, execution, review, and merge workflow nodes through authorized permanent principals with capacity leasing and recovery.
- Retire ephemeral workflow-stage workers and expose role-aware agent configuration, workflow editing, and documentation.
- Preserve lifecycle-column ratchet coverage by centralizing workflow-role classification rather than adding test exemptions.

Files changed:
 .changeset/fn-8764-workflow-role-agents.md         |   7 +
 CONCEPTS.md                                        |   3 +
 docs/agents.md                                     |   6 +
 docs/architecture.md                               |   6 +
 docs/cli-reference.md                              |   2 +
 docs/dashboard-guide.md                            |   4 +
 docs/settings-reference.md                         |   6 +-
 docs/storage.md                                    |   2 +
 docs/workflow-steps.md                             |   6 +
 .../src/__tests__/extension-agent-update.test.ts   |  11 +-
 packages/cli/src/__tests__/extension.test.ts       |  18 +-
 packages/cli/src/extension.ts                      |  41 +-
 .../core/src/__tests__/agent-permissions.test.ts   |  12 +
 .../core/src/__tests__/agent-role-policy.test.ts   |   7 +
 packages/core/src/__tests__/agent-roles.test.ts    |  21 +
 .../legacy-column-collection-gating-ledger.test.ts |  19 +-
 .../src/__tests__/postgres/schema-applier.test.ts  |  16 +-
 .../core/src/__tests__/settings-parity.test.ts     |   9 +-
 .../workflow-agent-node-classification.test.ts     |  25 +
 .../src/__tests__/workflow-work-item-cas.test.ts   |  38 ++
 packages/core/src/agents/agent-permissions.ts      |  11 +-
 packages/core/src/agents/agent-role-policy.ts      |  39 +-
 packages/core/src/agents/agent-store.ts            | 190 ++++++-
 .../core/src/async-stores/async-agent-store.ts     |   6 +
 packages/core/src/config/settings-schema.ts        |   5 +-
 packages/core/src/index.gate.ts                    |   2 +-
 packages/core/src/index.ts                         |   7 +-
 .../0045_fn_8764_multi_role_workflow_agents.sql    |  20 +
 .../0046_fn_8764_workflow_principal_fence.sql      |  49 ++
 packages/core/src/postgres/schema-applier.ts       |  22 +-
 packages/core/src/postgres/schema/project.ts       |  21 +
 packages/core/src/store.ts                         |   2 +-
 .../task-store/async/async-workflow-workitems.ts   |  49 +-
 packages/core/src/task-store/row-types.ts          |   4 +
 packages/core/src/task-store/settings-helpers.ts   |  16 +-
 packages/core/src/task-store/settings-ops-2.ts     |  13 +-
 packages/core/src/task-store/settings-ops.ts       |  16 +-
 packages/core/src/task-store/task-row-mappers.ts   |   4 +
 .../src/task-store/workflow-task-create-ops.ts     |   6 +-
 .../src/task-store/workflow-workitems-ops-2.ts     |  25 +-
 packages/core/src/types.ts                         |   2 +
 packages/core/src/types/agents/agents.ts           |  45 +-
 packages/core/src/types/merge/merge-queue.ts       |  17 +
 packages/core/src/types/settings/settings-scope.ts |   9 +-
 packages/core/src/workflows/workflow-ir-types.ts   |  58 +++
 packages/core/src/workflows/workflow-ir.ts         |  19 +
 .../dashboard/app/components/AgentDetailView.css   |  14 +
 .../dashboard/app/components/AgentDetailView.tsx   |  34 +-
 .../dashboard/app/components/NewAgentDialog.tsx    |  28 +-
 .../app/components/WorkflowNodeEditor.tsx          |  19 +
 .../__tests__/AgentDetailView.core.test.tsx        |   4 +-
 .../app/components/__tests__/AgentsView.test.tsx   |   2 +-
 .../__tests__/SettingsModal.general.test.tsx       |  86 ---
 .../__tests__/SettingsModal.test-harness.tsx       |   1 -
 .../components/agent-presets/agentCreatePayload.ts |   9 +-
 .../app/components/settings/section-keys.ts        |   1 -
 .../settings/sections/GeneralSection.tsx           |   8 -
 .../settings-default-descriptions.test.tsx         |   1 -
 .../app/components/workflow-flow-mapping.ts        |   7 +
 packages/dashboard/src/mission-routes.ts           |  26 +-
 .../src/routes/__tests__/agent-core-routes.test.ts |  23 +-
 .../src/routes/register-agent-core-routes.ts       |  42 +-
 ...gister-agent-import-export-generation-routes.ts |  21 -
 .../engine/src/__tests__/agent-action-gate.test.ts |  33 ++
 .../engine/src/__tests__/agent-assignment.test.ts  | 370 -------------
 .../src/__tests__/ephemeral-worker-manager.test.ts | 575 ---------------------
 ...ecutor-ephemeral-disabled-dispatch-gate.test.ts | 223 --------
 .../__tests__/executor-fast-mode-workflows.test.ts |  58 +++
 .../engine/src/__tests__/log-severity-manifest.ts  |   1 -
 .../__tests__/log-severity-spam-contract.test.ts   |   3 -
 .../resolved-read-with-literal-filter.test.ts      |   4 -
 .../__tests__/scheduler-ephemeral-toggle.test.ts   | 175 -------
 .../__tests__/scheduler-workflow-cutover.test.ts   |  19 -
 .../src/__tests__/workflow-agent-capacity.test.ts  |  47 ++
 .../src/__tests__/workflow-agent-routing.test.ts   | 137 +++++
 .../src/__tests__/workflow-graph-foreach.test.ts   |  15 +
 .../__tests__/workflow-graph-task-runner.test.ts   |  73 +++
 .../src/__tests__/workflow-task-runtime.test.ts    |  95 ++++
 .../src/__tests__/workflow-work-scheduler.test.ts  |  20 +
 packages/engine/src/agents/agent-action-gate.ts    |  64 +++
 packages/engine/src/agents/agent-assignment.ts     | 135 -----
 packages/engine/src/agents/agent-reflection.ts     |   1 +
 .../engine/src/agents/ephemeral-worker-manager.ts  | 429 ---------------
 .../engine/src/agents/workflow-agent-capacity.ts   | 113 ++++
 .../engine/src/agents/workflow-agent-router.ts     | 185 +++++++
 packages/engine/src/execution/reviewer.ts          |  26 +-
 packages/engine/src/executor.ts                    | 501 +++++++++++++++---
 packages/engine/src/index.ts                       |   1 -
 packages/engine/src/merger.ts                      |  20 +-
 packages/engine/src/pi.ts                          |  11 +
 packages/engine/src/runtimes/in-process-runtime.ts |  37 --
 packages/engine/src/scheduler.ts                   | 114 +---
 packages/engine/src/triage.ts                      | 196 ++++++-
 .../src/workflows/workflow-graph-executor.ts       | 109 +++-
 .../engine/src/workflows/workflow-graph-loop.ts    |  13 +-
 .../src/workflows/workflow-graph-task-runner.ts    |  12 +
 .../engine/src/workflows/workflow-task-runtime.ts  | 125 ++++-
 .../src/workflows/workflow-work-scheduler.ts       |   8 +-
 98 files changed, 2722 insertions(+), 2468 deletions(-)

Fusion-Task-Id: FN-8764
Fusion-Task-Lineage: 5527fccb-342d-46f6-8108-bbf89142efec
Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
2026-08-07 01:36:34 -07:00
gsxdsm
6bacfd74f5 fix(FN-8817): stop no-op lifecycle bounce
Honor verified intentional no-ops and preserve durable merger parks during workflow graph unwind.

Fusion-Task-Id: FN-8817
2026-08-06 16:11:05 -07:00
gsxdsm
4f4aef7173 FN-8811: preserve explicit shared-member review holds
Keep shared branch-group integration moving unless an operator explicitly holds the task.

- Track auto-merge provenance and distinguish explicit user holds from inherited mission policy.
- Preserve manual holds across workflow recovery, merge coordination, API updates, and dashboard status.
- Add regression coverage, document the behavior, and quarantine the observed flaky test.

Files changed:
 .changeset/fn-8811-shared-member-review-hold.md    |   7 ++
 docs/architecture.md                               |   4 +-
 docs/dashboard-guide.md                            |   1 +
 .../mission-store.sync-auto-merge.test.ts          |   7 +-
 .../__tests__/postgres/mission-store.pg.test.ts    |   1 +
 .../__tests__/postgres/store-movement.pg.test.ts   |  20 ++++
 packages/core/src/__tests__/task-merge.test.ts     |  14 +++
 .../core/src/async-stores/async-mission-store.ts   |   6 +-
 packages/core/src/index.gate.ts                    |   1 +
 packages/core/src/index.ts                         |   1 +
 packages/core/src/merge/task-merge.ts              |  20 +++-
 packages/core/src/missions/mission-store.ts        |   6 +-
 packages/core/src/task-store/serialization.ts      |   2 +-
 packages/core/src/task-store/task-creation.ts      |   8 +-
 packages/core/src/types/task/task-core.ts          |  12 ++-
 .../components/__tests__/TaskDetailModal.test.tsx  |  63 ++++++++++++
 .../dashboard/src/__tests__/routes-tasks.test.ts   |  47 +++++++++
 .../src/routes/register-task-workflow-routes.ts    |  15 ++-
 ...cutor-live-branch-group-auto-merge-hold.test.ts |  87 +++++++++++++++++
 .../src/__tests__/group-merge-coordinator.test.ts  |  99 ++++++++++++++++++-
 .../engine/src/__tests__/project-engine.test.ts    |  57 ++++++++++-
 .../self-healing-paused-abort-recovery.test.ts     |  52 +++++++++-
 packages/engine/src/__tests__/self-healing.test.ts | 106 +++++++++++++++++++++
 .../workflow-graph-executor-handlers.test.ts       |  23 +++++
 packages/engine/src/executor.ts                    |  37 ++++++-
 packages/engine/src/project-engine.ts              |  25 +++--
 packages/engine/src/self-healing.ts                |  71 ++++++++++++--
 .../src/workflow-node-runners/merge-runner.ts      |  24 ++++-
 .../src/workflows/workflow-graph-executor.ts       |   4 +
 .../src/workflows/workflow-graph-task-runner.ts    |   6 ++
 .../engine/src/workflows/workflow-node-handlers.ts |   5 +-
 packages/engine/vitest.config.ts                   |  11 ++-
 scripts/lib/test-quarantine.json                   |   5 +
 33 files changed, 789 insertions(+), 58 deletions(-)

Fusion-Task-Id: FN-8811

Fusion-Task-Lineage: 5c1609bf-3132-4988-a254-fedec6c0e33d

Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
2026-08-05 17:39:15 -07:00
gsxdsm
ec10411c4c FN-8795: persist structured workflow review findings
Persist normalized actionable findings from review workflow nodes.

- Normalize bounded finding IDs, text, locations, and severities in workflow results.
- Surface individual findings for Review-tab selection and same-task revision.
- Preserve findings through workflow retries and document the advisory contract.

Files changed:
 .changeset/fn-8795-structured-review-findings.md   |  7 +++
 docs/dashboard-guide.md                            |  2 +-
 docs/workflow-steps.md                             |  4 +-
 .../src/__tests__/workflow-step-results.test.ts    | 32 +++++++++++++-
 packages/core/src/index.gate.ts                    |  4 ++
 packages/core/src/index.ts                         |  6 ++-
 packages/core/src/types.ts                         |  4 ++
 packages/core/src/types/task/task-review.ts        |  5 +++
 packages/core/src/types/workflow/workflow-steps.ts | 20 +++++++++
 .../core/src/workflows/workflow-step-results.ts    | 50 +++++++++++++++++++++-
 .../dashboard/app/components/TaskReviewTab.tsx     | 12 +++++-
 .../src/routes/register-task-workflow-routes.ts    | 26 ++++++++++-
 .../workflow-malformed-verdict-gate.test.ts        | 12 ++++++
 packages/engine/src/executor.ts                    | 46 +++++++++++++++++---
 .../src/workflows/workflow-graph-executor.ts       | 11 +++++
 15 files changed, 227 insertions(+), 14 deletions(-)

Fusion-Task-Id: FN-8795
Fusion-Task-Lineage: 09003b01-3f9a-4387-b6a7-f29066ce52f6
Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
2026-08-05 00:00:03 -07:00
gsxdsm
1dc636b8b4 FN-8785: deduplicate queued dependency and scope logs
Persist queue episodes atomically so repeated scheduler and self-healing passes do not duplicate diagnostics.

- Add a queued-episode signature with PostgreSQL migration and task serialization support.
- Route dependency and file-scope queue transitions through the atomic deduplication API.
- Cover repeated and concurrent queue transitions, and update scheduler mocks for the new store API.

Files changed:
 .changeset/fn-8785-queued-log-deduplication.md     |   7 +
 docs/architecture.md                               |   1 +
 .../postgres/queued-episode-transition.pg.test.ts  | 158 +++++++++++++++++++++
 .../src/__tests__/postgres/schema-applier.test.ts  |   9 +-
 .../core/src/postgres/migrations/0000_initial.sql  |   1 +
 .../0044_fn_8785_queued_episode_signature.sql      |   3 +
 packages/core/src/postgres/schema-applier.ts       |  12 +-
 packages/core/src/postgres/schema/project.ts       |   1 +
 packages/core/src/store.ts                         |   5 +-
 packages/core/src/task-store/audit-ops.ts          |  80 +++++++++++
 packages/core/src/task-store/persistence.ts        |   2 +
 packages/core/src/task-store/serialization.ts      |   1 +
 packages/core/src/types/task/task-core.ts          |   5 +
 ...executor-outer-dispatch-dependency-gate.test.ts |  39 +++--
 .../engine/src/__tests__/executor-test-helpers.ts  |   7 +
 .../__tests__/scheduler-overlap-starvation.test.ts |  43 +++++-
 .../__tests__/scheduler-workflow-cutover.test.ts   |  35 +++--
 .../self-healing-completion-fanout.test.ts         |  37 +++++
 packages/engine/src/__tests__/self-healing.test.ts |  25 +++-
 packages/engine/src/executor.ts                    |  16 ++-
 packages/engine/src/scheduler.ts                   |  26 ++--
 packages/engine/src/self-healing.ts                | 117 +++++----------
 22 files changed, 491 insertions(+), 139 deletions(-)

Fusion-Task-Id: FN-8785
Fusion-Task-Lineage: 8d68c243-f24e-4db4-bd1d-b7eda01309f6
Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
2026-08-04 11:52:15 -07:00
gsxdsm
9939897aab fix: prevent stale planning approvals and review churn (#3327)
## Summary

Planning can no longer approve or execute against evidence from a
superseded dependency episode. Dependency mutations, approval decisions,
recovery, and execution admission now share serialized lifecycle rules,
so stale planner work cannot restore an invalid approval or release an
unplanned task.

Review also converges instead of discovering one blocker per round.
Planning performs a repository-grounded completeness pass up front; Plan
Review batches all independently discoverable blockers and carries an
episode-scoped decision ledger across revisions; code review traces
changed invariants through production consumers and tests. Repeated
feedback still advances the safety budget, while provider failures and
superseded episodes stay outside the remediation ledger.

The dashboard now exposes manual approval only for the intended
exhausted-review state, and refusal/recovery audit events make rejected
lifecycle transitions diagnosable without leaking prompt content.

## Validation

- `pnpm verify:fast` — scoped typechecks/builds, CLI build, and boot
smoke passed.
- Focused Core and Engine regression suites — 511 tests passed.
- `pnpm lint`, strict changeset validation, Core/Engine typechecks, and
package builds passed.

Fixes #3325.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Improved Plan Review approvals, rejections, and replan-cap handling
across task workflows.
* Added cumulative feedback and attempt tracking across repeated
planning reviews.
* Added safer recovery for stalled planning handoffs and interrupted
approval updates.
* **Bug Fixes**
* Prevented stale approvals and unplanned execution after dependency
changes.
* Improved concurrent approval handling, retryability, and
refusal-record deduplication.
  * Refined dashboard approval indicators and responsive approval views.
* **Quality Improvements**
* Strengthened planning and code-review completeness checks and
blocking-finding coverage.
  * Preserved review history while clearly marking outdated approvals.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-08-04 00:20:08 -07:00
gsxdsm
7824150715 FN-8769: protect default-branch mission merges
Keep mission group members behind manual release controls when their target is the default branch.

- Create deterministic intermediate branches for project-default mission groups.
- Gate default-branch group routing and auto-merge exemptions behind the normal manual-release flow.
- Cover intermediate and default-branch group behavior with core and engine tests.

Files changed:
 ...fn-8769-default-branch-group-auto-merge-gate.md |   7 ++
 docs/architecture.md                               |   2 +-
 docs/missions.md                                   |   2 +-
 .../mission-store.sync-auto-merge.test.ts          |   4 +-
 packages/core/src/__tests__/task-merge.test.ts     |  16 ++-
 .../core/src/async-stores/async-mission-store.ts   |  12 ++-
 packages/core/src/merge/task-merge.ts              |  23 +++-
 packages/core/src/missions/mission-store.ts        |  11 +-
 ...cutor-live-branch-group-auto-merge-hold.test.ts |  16 ++-
 .../src/__tests__/group-merge-coordinator.test.ts  | 118 ++++++++++++++++++++-
 .../engine/src/__tests__/project-engine.test.ts    |  16 ++-
 packages/engine/src/executor.ts                    |   4 +-
 .../engine/src/merge/group-merge-coordinator.ts    |  13 +++
 packages/engine/src/project-engine.ts              |  15 ++-
 14 files changed, 235 insertions(+), 24 deletions(-)

Fusion-Task-Id: FN-8769

Fusion-Task-Lineage: dff96c8e-ca96-437c-94bf-9691bdf572e9

Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
2026-08-03 16:48:25 -07:00
gsxdsm
4ff41a723c fix: self-heal executor credential resolution for custom providers
Stop synthesizing credentialInstanceId "default" into executor sessions,
soft-fail unresolved instances to the legacy unscoped auth path, and
collapse-match renamed custom-provider auth slugs so task execute matches chat.
2026-08-03 10:58:19 -07:00
gsxdsm
cb57093d03 refactor: domain folder layout (types, API, core, engine) (#2398)
## Summary

Wave 17 organizes Fusion into **domain folders** (stacks on #2397).

### Layout
- **core/types/** — board, task, agents, settings, merge, workflow,
mesh, …
- **core/src/** — agents, ai, async-stores, workflows, tasks, config,
db, …
- **dashboard/app/api/** — client, tasks, agents, git, missions,
planning, …
- **engine/src/** — agents, auth, execution, merge, missions, overseer,
worktree, …

Root keepers retained for large entrypoints (`store.ts`, `executor.ts`,
`merger.ts`, …).

Public barrels (`@fusion/core`, `@fusion/engine`, `app/api.ts` → legacy)
stay stable.

## Test plan
- [x] `@fusion/core` typecheck
- [x] `@fusion/engine` typecheck (pre-existing playwright-core noise
only)
- [ ] CI merge gate

**Stack:** #2394 → #2397 → **this PR**
2026-08-03 00:20:53 -07:00
gsxdsm
73fe461db5 fix: demote routine engine log noise to debug
Keep the default TUI for lifecycle transitions (Starting/Specifying/Worktree created/merge/move). Demote expected skips, schedule-trigger echoes, session setup bookkeeping, warm worktree reuse, and fn_run_verification command-fail detail. Pin via severity manifest and contract tests.
2026-08-02 23:18:07 -07:00
gsxdsm
1e7f510ee2 fix: stop blocking tasks on open-PR file claims — board tasks are the only blockers
Remove the FN-8700 PR/file-claim blocking mechanism end to end (operator
decision after FN-8728 parked on unrelated PR #2398):

- Drop the AGENTS.md claim-check rule and scripts/check-file-claimed.mjs
- Executor prompt + fn_task_done no longer accept pr:N refs or treat open
  PRs as blocked-exit reasons
- execution-block-classifier classifies on Fusion task dependencies only;
  legacy pr refs are discarded, reason prose never makes a block durable
- Remove the session-log BLOCKED promotion and the gh-backed
  reconcile-external-pr-blockers self-healing sweep
- Legacy file-claim parks are no longer honored, so previously PR-blocked
  rows recover via normal paths

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-02 17:06:38 -07:00
gsxdsm
9e0f29abc9 fix: durable parks for file-claim and PR blockers (FN-8700)
Classify blocked exits so check-file-claimed / open-PR collisions never
auto-replan. Park failed with externalBlockers metadata (pr:N supported),
promote BLOCKED session logs instead of incomplete-step requeue, thrash-
exhaust after 3 identical durable blocks, and clear parks when gh reports
blocking PRs merged or closed.
2026-08-01 18:34:42 -07:00
gsxdsm
6e98e16e70 fix: reject DUPLICATE-only PROMPT at dispatch (FN-8704)
FN-8704 failed at the graph parse node because PROMPT.md was only
"DUPLICATE: FN-8676". Filesystem validation treated non-empty as planned
and admitted the card into WIP, which then looped on parse failure.

Treat a sole DUPLICATE redirect as unplanned: block dispatch and hold
release, badge as awaiting planning, and if parse still sees that shape
rebound to needs-replan with feedback instead of parking failed.
2026-08-01 12:26:53 -07:00
gsxdsm
60706ed5e4 fix: demote high-frequency TUI log lines to debug
Session setup, track bookkeeping, intentional skill exclusions, token-cache
metrics, zero-count recovery summaries, and expected-missing PROMPT seed reads
were flooding the default log pane. Gate them behind FUSION_DEBUG so only
state transitions and operator-actionable warnings remain visible.
2026-08-01 11:48:56 -07:00
gsxdsm
01d65805c3 FN-8693: refresh reused worktree bases before execution
Refresh reused execution worktrees against the current integration baseline.

- Rebase or reset clean reused worktrees before coding sessions while preserving task commits.
- Persist and audit refreshed base SHAs, and block unsafe refresh states before execution.
- Cover executor, graph, and heartbeat refresh paths with regression tests.

Files changed:
 .changeset/fn-8693-stale-worktree-base.md          |   7 +
 docs/architecture.md                               |   1 +
 .../src/__tests__/agent-heartbeat-worktree.test.ts |  28 ++++
 .../__tests__/ce-workflow-step-executor.test.ts    |  44 ++++++
 .../src/__tests__/worktree-base-refresh.test.ts    |  90 ++++++++++++
 packages/engine/src/agent-heartbeat.ts             |  35 ++++-
 packages/engine/src/executor.ts                    |  67 ++++++++-
 packages/engine/src/merger.ts                      |   5 +
 packages/engine/src/run-audit.ts                   |  11 ++
 packages/engine/src/workflow-graph-executor.ts     |  32 ++++-
 packages/engine/src/worktree-acquisition.ts        |  28 +++-
 packages/engine/src/worktree-base-refresh.ts       | 158 +++++++++++++++++++++
 12 files changed, 498 insertions(+), 8 deletions(-)

Fusion-Task-Id: FN-8693
Fusion-Task-Lineage: e39a441f-39b5-4723-b503-753e921018f3
Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
2026-08-01 10:03:28 -07:00
gsxdsm
04c2bb4707 FN-8654: rotate credential instances after provider limits
Retry provider-limit failures with eligible credential instances before falling back to existing pauses and backoff.

- Add a runtime-shared credential rotator with cooldown, exhaustion, and audit handling.
- Wire credential rotation into executor and heartbeat retry lanes while preserving user pause controls.
- Document the behavior and cover rotation, recovery, and retry paths.

Files changed:
 .changeset/fn-8654-credential-instance-rotation.md |   7 +
 AGENTS.md                                          |   1 +
 docs/architecture.md                               |   2 +-
 docs/settings-reference.md                         |   4 +
 .../__tests__/credential-instance-rotation.test.ts |  88 +++++++++++
 .../__tests__/credential-rotation-lanes.test.ts    |  20 +++
 .../__tests__/credential-rotation-recovery.test.ts |  19 +++
 .../__tests__/credential-rotation-wiring.test.ts   |  15 ++
 .../__tests__/rate-limit-retry-rotation.test.ts    |  50 ++++++
 .../src/__tests__/usage-limit-detector.test.ts     |  14 ++
 packages/engine/src/agent-heartbeat.ts             | 102 +++++++++++-
 .../engine/src/credential-instance-rotation.ts     | 175 +++++++++++++++++++++
 packages/engine/src/executor.ts                    | 141 +++++++++++++++--
 packages/engine/src/index.ts                       |   7 +
 packages/engine/src/project-engine.ts              |   5 +
 packages/engine/src/rate-limit-retry.ts            |  32 +++-
 packages/engine/src/runtimes/in-process-runtime.ts |  29 +++-
 packages/engine/src/usage-limit-detector.ts        |  18 ++-
 18 files changed, 699 insertions(+), 30 deletions(-)

Fusion-Task-Id: FN-8654
Fusion-Task-Lineage: 44d63441-270c-4949-8c34-47ec4c9992e4
Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
2026-08-01 08:20:29 -07:00
gsxdsm
4f09758dc8 FN-8681: retarget executor step credential instances
Enable executor step sessions to use selected and rotated credential instances.

- Pass task-selected credential instances into step-session execution.
- Re-resolve live credential targets after usage-limit retries using the effective agent runtime configuration.
- Retarget future sessions safely and cover retry behavior.
- Document the runtime behavior and add a patch changeset.

Files changed:
 .changeset/fn-8681-credential-instance-retarget.md |   7 +
 docs/settings-reference.md                         |   7 +-
 .../src/__tests__/step-session-executor.test.ts    | 211 +++++++++++++++++++++
 packages/engine/src/executor.ts                    |  21 ++
 packages/engine/src/step-session-executor.ts       |  94 ++++++++-
 5 files changed, 332 insertions(+), 8 deletions(-)

Fusion-Task-Id: FN-8681

Fusion-Task-Lineage: 99724283-6f5d-4890-b7f8-65af1d88b12c

Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
2026-08-01 03:45:59 -07:00
gsxdsm
8a6949dd24 FN-8661: resolve selected credential instances for sessions
Resolve requested provider credential instances before creating agent sessions.

- Thread lane credential instance selections through planning, validation, execution, review, and merge sessions.
- Resolve selected instances into runtime credential stores while retaining provider-default fallback behavior.
- Preserve selected credentials for mission validation, executor retries, and spawned child agents.

Files changed:
 .../fn-8661-credential-instance-resolution.md      |  7 ++
 AGENTS.md                                          |  1 +
 docs/architecture.md                               |  2 +-
 docs/secrets.md                                    |  2 +
 docs/settings-reference.md                         |  1 +
 .../dashboard/src/__tests__/routes-auth.test.ts    | 80 +++++++++++++++++++
 .../dashboard/src/routes/register-model-routes.ts  | 78 +++++++++++++++++++
 .../src/__tests__/agent-session-helpers.test.ts    | 16 ++++
 .../credential-instance-resolution.test.ts         | 49 ++++++++++++
 packages/engine/src/agent-heartbeat.ts             |  1 +
 packages/engine/src/agent-runtime.ts               |  9 ++-
 packages/engine/src/agent-session-helpers.ts       | 67 +++++++++++-----
 packages/engine/src/auth-storage.ts                | 90 ++++++++++++++++++----
 packages/engine/src/executor.ts                    | 29 ++++++-
 packages/engine/src/merger-ai.ts                   |  2 +
 packages/engine/src/merger.ts                      |  5 ++
 packages/engine/src/mission-execution-loop.ts      |  4 +-
 packages/engine/src/pi.ts                          |  7 +-
 packages/engine/src/pr-response-run-ops.ts         |  1 +
 packages/engine/src/reviewer.ts                    |  7 ++
 packages/engine/src/triage.ts                      |  2 +
 21 files changed, 420 insertions(+), 40 deletions(-)

Fusion-Task-Id: FN-8661

Fusion-Task-Lineage: 1e34a3ce-0857-4619-9746-ce0dc12dc2ba

Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
2026-08-01 02:05:03 -07:00
gsxdsm
5a19d1da6e fix: count only actively running tasks against worktree capacity
Retained directories on queued, paused, blocked, or terminal tasks no longer
consume scheduler slots. Agent concurrency and worktree capacity now count the
same canonical live-task population through one project admission ceiling
(resolveActiveTaskCapacityLimit) with an atomic reserveIfAvailable claim, so
planning, execute, and merge lanes cannot each observe and claim the final
worktree slot independently.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-31 22:07:05 -07:00
gsxdsm
7bdeaa8b0c fix(engine): the new worktree ledgers count terminal lanes by NAME — renamed boards stall (#3296)
## Census

**Before: `COLUMN guards (the backlog): 2`, `--strict` RED. After:
`BACKLOG ZERO`, all five gates green.**

Two commits from last night's `maxWorktrees` rollout copied the same
holder ledger, both with literals:

| commit | file | gate |
|---|---|---|
| `374956ef23` | `triage.ts` | planning admission |
| `6c7467a78d` | `executor.ts` | `fn_spawn_agent` |

```ts
t.column !== "done" && t.column !== "archived"
```

## What it costs

Both exclude terminal lanes because a finished card's worktree is
**cleanup-owned, not capacity**. On a renamed board neither literal
matches, so every finished card keeps counting as a live holder. The
count only grows, the gate reaches zero room on a board with free slots,
and planning admission is withheld forever / every spawn is refused.

That is the **mirror** of the breach these commits fixed, and strictly
worse: 8 planners on a 4-slot board is visible; a permanent stall is
silent. The recorded reason even names the worktree budget, which the
operator then checks and finds has room.

## The conversion

`resolveProjectColumnsForRoles(store, ["complete", "archived"])` —
project-level, because the ledger spans the whole board with no single
task to resolve against. Matches triage's existing use in
`sweepStalePlanningStatuses` and executor's at the wip gates.
Legacy-seeded, so a default board still excludes exactly `done` and
`archived` — byte-identical there.

## Both conversions were UNCOVERED when written

Measured with #3214's blinding procedure **before** writing tests:
reverting either to the literals left **all 19 tests in the capacity
suites green**. Nothing in the tree could tell the conversion from what
it replaced — which is how the literals got there in the first place.

Each now has a renamed-board case that fails when blinded:

```
triage    converted 2 passed  |  BLINDED 1 failed | 1 passed  |  restored 2 passed
executor  converted 8 passed  |  BLINDED 1 failed | 7 passed  |  restored 8 passed
```

## The pairing earned itself immediately

Both new cases assert an **absence** (no throttle / no refusal), so each
is paired with a positive proving the gate still fires on the same
renamed board when a card genuinely holds the last worktree.

That caught a real defect in my own fixture: the candidate scan resolves
each task's **own workflow selection**, not `listWorkflowDefinitions`,
so my first version fell back to the default board where `drafting`
isn't a hold lane. No card was eligible, nothing throttled, and the
absence assertion **passed for the wrong reason**. The positive failed
and exposed it. Recorded at the fixture so the next reader doesn't
reintroduce it.

## Verification

```
42 tests across 6 capacity suites             pass
check-fnxc-future-dates                       green
check-inert-sync-lane-conversions             green
check-lane-wiring                             green
check-sql-column-literals                     green
census --strict                               green   (BACKLOG ZERO restored)
```

No changeset: internal engine fix, no published-package surface change.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 19:16:07 -07:00
gsxdsm
6c7467a78d fix(engine): spawned children gate on maxWorktrees too — the spawn note promised both dimensions
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-31 18:43:37 -07:00
gsxdsm
500f40e65b fix: descriptive waiting badges (Queued to revise / Queued behind FN-X) + dependency-free blocked exits replan calmly
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-31 17:37:38 -07:00
gsxdsm
10a0c5848f fix(executor): planner-evacuation lanes come from the emitter — executor leaves the inert list (16 → 12) (#3137)
`executor.ts` was the last file besides `triage.ts` and `scheduler.ts`
on `check-inert-sync-lanes`, holding **4 guards that read as converted
and behave as literals**. Neither cause turned out to be "needs an async
resolver".

## 1. Two of the four were in code with no caller

`isPlannerColumnFor` is a **private method with zero production
callers**. `tsc` reports it unused; the only things reaching it were two
tests casting through `executor as unknown as { … }`, which is exactly
what let it look alive. Its doc comment described the
planning-evacuation branch — but that branch calls
`isBackwardMoveOutOfPlanning` and never called this.

Deleted, along with the two tests whose subject it was. Converting
guards in unreachable code would have "fixed" behaviour that cannot run
and left two more sites to maintain; a test whose subject has no caller
pins nothing.

## 2. The other two no longer need to resolve anything

`isBackwardMoveOutOfPlanning` resolved its own lanes via
`resolvePlannerLanes`, whose selection reader returns `undefined`
unconditionally under PostgreSQL — so it answered with the **default
board for every task**, and both its guards were inert.

Its comment justified the sync resolver by the synchronous `task:moved`
emitter. That was true and **is no longer binding**: the emitter now
resolves lanes once, asynchronously (`moves.ts` →
`resolveWorkflowIrForTask`), and hands them on the payload — which #3112
already reads in this same listener. Reading a parameter is as
synchronous as reading `from`, so nothing reorders and no listener
resolves.

`lanes` is **required, not optional**. An optional parameter that the
one production caller happens to pass is the seam-with-no-supplier shape
this program keeps finding; required means a future caller fails
typecheck instead of silently getting a default board. When the emitter
itself could not resolve, the legacy ids answer — exactly what
`resolvePlannerLanes` degraded to anyway.

## Measured

| | before | after |
|---|---|---|
| `check-inert-sync-lanes` | **16** guards, 3 files | **12** guards, 2
files |
| `executor.ts` on that list | 4 | **0 — off the list** |
| census | 18 | 18 (`--strict`: every file matches baseline exactly) |

**The census is deliberately unchanged.** This targets the inert
population, which the census cannot see by construction: those guards
already read as converted. That gap is the argument in #3082 — 12 guards
still behave as literals while the census shows them as done.

## The producer half, which I nearly shipped without

The predicate's own suite covers it thoroughly — and every case calls it
**directly**. Mutation testing exposed that this proves nothing about
the listener: replacing the listener's `lanes` argument with `undefined`
left `planning-evacuation` at **20/20 green**. That is the fifth failure
shape in this program's learnings verbatim — a converted consumer with
an unconverted producer passing every instrument.

So there is now a case driving the **real listener** on a board whose
planner lanes share no id with the legacy pair (`queued` holds,
`drafting` intakes), withdrawing a card to a non-lifecycle column — the
reported symptom (`todo -> Ideas`) in that board's vocabulary.

## Verification

- engine `tsc` — **0 errors**
- `executor-planner-lanes-resolved` — **12 passed**
- `executor-archive-releases-active-session` — **14 passed**; listener
passing `undefined` → **1 failed | 13 passed**
- `planning-evacuation` + `triage-planning-wake` + archive suite — **47
passed**
- `check-inert-flag-seams`, `check-fnxc-future-dates`, census `--strict`
— exit 0
- `eslint` on changed files — 0 errors

The predicate tests are also **stronger than before**, not merely
adapted: they now build lanes with `toTaskMoveLanes`, the same function
`moves.ts` uses for the payload. Previously they reached the predicate
through the store-backed sync reader, so renamed-lane assertions passed
in the harness while the real path could never see a renamed lane.

## Not done here

The inert baseline still reads 29 against a tree of 12 and the gate
advises re-recording. I left it: a stale allowance is a real hazard, but
re-recording is a one-line change that conflicts with every lane, and it
should land once rather than in each of our branches.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 07:30:43 -07:00
gsxdsm
4b61170a51 fix(executor): read task:moved lanes from the payload (executor.ts 4 → 0) (#3112)
**Stacked on #3109** — merge that first; this is its first consumer.

## Census

| Metric | Before | After |
|---|---:|---:|
| COLUMN guards (backlog) | 47 | **43** |
| `executor.ts` | 4 | **0** |

`executor.ts` is off the census top-files list.

## Why these four could not be converted in place

This listener is synchronous and its branches **start execution**,
dispose worktrees and release sessions. An await ahead of them defers
the `execute()` dispatch itself. The sync IR resolver isn't an option
either — it answers with the default workflow under PostgreSQL, so a
guard written through it is inert.

Reading the lanes the emitter already resolved costs nothing and leaves
the prologue synchronous. This listener is the reason #3109 has the
shape it does.

## The archive branch is the one with teeth

`to === "archived"` matched nothing on a board with a renamed terminal
lane, so **archiving never released the task's active-session registry
entry** — and that entry is what blocks a **successor** task from
acquiring the same path. Not cosmetic: the next task wanting that path
fails to register.

## Verification

- **Revert-proof:** the new case drives a `shipped` terminal lane
(matching no legacy id) and asserts the release. Reverting the branch to
the literal leaves the entry held — `expected [Array(1)] to have a
length of 0`.
- 43 executor suites — **483 green**
- **`pnpm test:gate` green**; eslint clean

## Note on shape

Lanes are read as **single ids, not sets**, because each branch here is
a lane-identity test on one column — exactly what the literals were.
Widening to membership would change behaviour, not just vocabulary.
Fail-soft to the legacy ids when the emit path could not resolve,
matching every other consumer of this payload.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 05:15:38 -07:00
gsxdsm
15a664a8f5 docs(engine): flag executor's four task:moved literals — the obvious conversion is provably inert (#3104)
The largest unclaimed census cluster. **Nothing in this file said why
the sync-lane pass skipped it**, and that silence is the hazard: the
obvious next move is to convert these the way `scheduler.ts`'s ten were
converted, which would make them **inert rather than fixed**.

## The literals are genuinely wrong — this is not a "non-issue" flag

All four sit in one synchronous `task:moved` listener, and on a renamed
board:

- execution **never starts** on a move into the board's own wip lane;
- terminal session release **never runs** on a move into its archive
lane;
- both `from` guards never fire, so **in-flight work is not aborted**
when a card leaves implementation.

Nothing errors. The engine simply stops reacting.

## Why the obvious fix is inert — proved, not argued

`task:moved` is emitted synchronously, so an `await` here reorders this
handler against every other subscriber. That points at the sync IR path,
which cannot answer for a renamed board for **two independent reasons**
(`sync-workflow-ir-second-blocker.test.ts`, #3103):

1. `getTaskWorkflowSelectionImpl` returns `undefined` unconditionally
under PostgreSQL, so `resolveTaskWorkflowIrSync` always takes its
`!workflowId` branch.
2. Even **with** a selection, the custom-workflow branch loads its IR
through `store.db`, whose implementation is an **unconditional throw** —
so it falls into the catch and returns the default IR anyway.

**A renamed lane is a custom workflow, so (2) alone is decisive.** The
sync path can never serve this listener's case, whatever the selection
reader is fixed to do. That is the part the existing notes across this
repo miss, and it is why flagging beats attempting here.

`check-inert-sync-lane-conversions` already baselines **twenty** guards
in exactly that state in `scheduler.ts`. These four must not join them.

## Census

**Unchanged at 4, deliberately.**

Marking them DELIBERATE-LITERAL would buy a smaller number by asserting
the code is *fine*. It is not fine — it is *blocked*. Those are
different claims with different expiries, and the census should keep
pointing here until the block is lifted. An unconverted literal is
visible; an inert conversion leaves the backlog and takes the evidence
with it.

## Measured

- Comment-only change.
- `src/__tests__/executor*` — **84 files / 853 tests pass**.
- `tsc --noEmit -p packages/engine` clean; census `--strict`,
`check-inert-sync-lane-conversions`, `check-fnxc-future-dates` clean.

## Unblocking, for whoever takes it

Either an async listener contract — a behaviour change to handler
ordering, not a column conversion — or a sync reader that answers for
**custom** workflows *and* survives a writer on another node. All three
constraints are written up in `sync-workflow-ir-second-blocker.test.ts`
(#3103).

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 04:40:42 -07:00
gsxdsm
eb0ee4ae98 fleet: executor.ts 7 → 4 lifecycle-column guards (3 converted, 4 flagged out of scope) (#3048)
Claiming `packages/engine/src/executor.ts` from the census work order.

## Census before/after

| File | Before | After |
|---|---:|---:|
| `packages/engine/src/executor.ts` | 7 | **4** |

Measured with `scripts/lifecycle-column-census.mjs` (kind `column`
only), not grep.

## Converted (3)

**L17258 — the completed-task watchdog never armed on a renamed board.**
It required the card to sit in a literal `in-progress`. This does not
error; the watchdog simply never fires, which is the silent-guard class
this program exists to remove. The branch immediately above already
resolves the same lane through `resolveWipTargetForTask`, and there is
even an FNXC note there saying `latestColumn` must come from that
resolved value — so the comparison now asks the same resolver rather
than an id.

**L14940 (×2) — the duplicate-handoff finalize never ran on a renamed
review lane.** `fromColumn`/`toColumn` are parsed out of the store's
rejection message (`Invalid transition: 'X' → 'Y'`), so they carry
whatever ids that workflow declares. Comparing them to the literal
`in-review` meant a renamed lane never matched and
`finalizeAlreadyReviewedTask` was skipped, leaving the card
mid-transition with nothing to complete it. Now resolves the task's own
review role, falling back to the legacy literal when the workflow cannot
be read — so behaviour is unchanged wherever the vocabulary is
unreadable.

## Flagged, not converted (4) — per the fleet rule that behavior changes
are out of scope

**L3557 / L3581 / L3632 / L3642** are branch conditions inside the
**synchronous** `store.on("task:moved")` listener. Resolving a task's
workflow requires an `await`, which is not available in a sync
listener's condition. Moving the test into the deferred body would widen
the branch to every non-forward move and then re-narrow it — a
**behaviour change to the planning-evacuation path**, not a vocabulary
conversion. Converting them properly means making the listener async,
which wants its own commit and its own test.

I flagged rather than guessed, which is why this is 7 → 4 and not 7 → 0.

## On test coverage, stated plainly

Both converted sites are pure resolver swaps in `async` contexts,
verified by tsc, the census delta, and the existing executor suites (48
tests green). I did **not** add new fixtures: this is the file where I
twice wrote tests that passed against the *unconverted* code —
`recoverCompletedTask`'s seven early-return guards make negative
assertions succeed trivially — and reverted both times rather than claim
coverage I did not have. A fixture that genuinely drives L17258 needs a
satisfied `workflowStepResults` so the run does not divert into graph
re-entry; that is worth doing, and it is worth doing honestly rather
than as a green-looking placeholder.

## Verification

- census: `executor.ts` 7 → 4
- `tsc --noEmit` on `@fusion/engine` clean; `pnpm lint` clean
- `executor-graph-boundary`, `executor-task-done-summary`,
`executor-triage-column-audit`, `executor-step-session` — 48 tests green

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 02:27:49 -07:00
gsxdsm
189f237a07 fix(security): make the approval and permission gates actually enforce (#2457)
## Why

An AI agent autonomously deleted a live task (FN-8609) via
`fn_task_delete`. Investigating that led to two audits, which found the
gate machinery is well-built and its *enforcement* well-tested — but the
*decision boundary* was unguarded and the tools that matter never
reached the gate at all.

The headline defect: all ~76 tools in `packages/cli/src/extension.ts`
are delivered through `hostExtensionPaths` → `DefaultResourceLoader`,
which never passes through `wrapToolsWithActionGate`. `fn_task_delete`
*was* classified as gated in `gating-classifications.ts` — that
classification was dead code. Roughly 15 destructive tools were ungated
even for an operator who had locked their policy all the way down.

Second: `POST /api/approvals/:id/decision` took the approver identity
from the request body, validated for shape only, with no check that the
decider was human or differed from the requester. Combined with a bash
tool whose command string was never inspected (only `params.cwd`), an
agent could read the daemon token from `~/.fusion/settings.json`, find
its own pending request, and approve it.

## Scope

This does **not** tighten defaults. The `unrestricted` preset is
byte-for-byte unchanged — `git diff` on `agent-permission-policy.ts` is
empty — and regression tests assert that an out-of-the-box install
behaves exactly as before. The bug was never "the default is too
permissive"; it was "strict policy doesn't enforce." This makes turning
security up actually work.

The one deliberate exception: the containment that stops an agent
escalating its *own* privileges (reading the daemon token / credentials,
calling the approvals API to self-approve) applies at every preset
including `unrestricted`. That is a privilege-escalation boundary rather
than a permission preference — if it only engaged under strict policy it
would not have prevented the incident that prompted this.

## What changed

8 bisectable commits:

- **Approval lifecycle** — self-approval blocked via server-derived
deciders; same-verdict replay 409s; decide re-reads and re-validates
inside the transaction; expiry TTLs; `markCompleted` ownership check;
session identity registry in core.
- **Engine gates enforce for real** — unclassified tools resolve to a
policy-governed category instead of hardcoded `allow`; missing-policy
fail-open closed; bash containment floor + exact-command approval
binding.
- **Dashboard decision routes** — stop trusting client-supplied actors
(decision, bypass-review, worktrunk → 403 on forged actors).
- **`fn serve` authenticated by default** — auto-mints a token following
the existing `fn dashboard` precedent; `--no-auth` opts out.
- **Sibling entry points closed** — user-sourced hard-cancel moves, ACP
execute-once approvals, plugin task-store gating.
- **pi-extension principal resolution** — the extension resolves the
acting principal and can withhold or policy-gate the previously ungated
destructive tools.
- **Root-cause bonus fix** — `findLatestByDedupeKey` was broken in
PostgreSQL backend mode (already-parsed jsonb fed through a string-only
parser), so approved-grant redemption **never matched in production**,
minting duplicate requests. This explains the live DB state of 17
approved / 0 completed. *(Also cherry-picked to `main` as `a9b30013bb`,
since it is an active production defect on its own.)*
- **Review follow-ups** (`627f1b1fa8`) — operator-configured
provisioning privilege and a configurable grant TTL; see below.

## Review follow-ups

**Provisioning privilege is operator-configured, not role-derived.**
`isCallerPrivileged` had gone from `caller.reportsTo == null` (every
top-level agent privileged — permanent escalation by creating a
manager-less agent) to `caller.role === "ceo"`, which swapped an
implicit rule for a magic string: any agent config can claim that role,
while an operator who genuinely wants a privileged agent had no
supported way to say so. Privilege now derives solely from
`agentProvisioning.trustedAgentIds` / `trustedRoles` and fails closed
when settings are unresolvable.

It is also no longer forwarded to `resolveAgentProvisioningPolicy` as
`isPrivileged`, because that flag short-circuits ahead of
`alwaysApproveDelete` — a trusted caller was bypassing delete approval
entirely. The policy applies the same trusted rules itself, in the right
order. The function now governs only the org-chart escape hatch (acting
outside your own direct reports).

**Grant TTL defaults to 1 hour and is configurable.** Approval →
redemption is not instantaneous: an operator approving from their phone,
an engine restart, a queued lane, or a task waiting on a worktree all
routinely exceeded 15 minutes, after which the grant expired and the
agent silently re-requested. One hour remains far short of the
"redeemable forever" hazard the TTL exists to bound. Override via
`FUSION_APPROVAL_GRANT_TTL_MS` or `configureApprovalRequestTtls()`;
invalid overrides are ignored rather than widening the window to
infinity or collapsing it to zero.

## Behavior changes requiring operator review before rollout

1. `fn serve` requires a bearer token by default (`--no-auth` opts out);
unauthenticated clients get 401.
2. Agents can no longer run withheld destructive tools
(`fn_task_delete`, `fn_task_bypass_review`,
mission/milestone/slice/feature/workflow deletes, `experiment_finalize`,
`skills_install`). Operators keep them via CLI/dashboard. **This is the
incident fix.**
3. Agents get provisioning privilege only when the operator lists them
in `agentProvisioning.trustedAgentIds` / `trustedRoles`; the
provisioning gate is now live in production. Previously-implicit
privilege (top-level position, or a `ceo` role) no longer grants
anything on its own.
4. Decision replay 409s (was 200); pending approvals expire after 24h,
approved grants after 1h (configurable); bash approvals bind per exact
command.
5. Forged/body actors on decision, bypass-review, worktrunk routes →
403; `archive-all-done` requires `{confirm:true}` (external scripts
affected).
6. `fn_secret_get` approvals grant exactly one reveal (previously
granted nothing and looped forever); ACP approvals are execute-once
(previously infinite reuse).
7. Bash containment denies token/credential/approvals-API commands in
all agent sessions at every preset.

## Verification

Independently re-run against the branch, not just self-reported:

- 5 typechecks (core, engine, cli, dashboard `tsconfig.json` +
`tsconfig.app.json`) — clean
- `pnpm lint` — clean
- `pnpm test:gate` — 379 passed
- `pnpm build --force` — green (a plain `pnpm build` skips packages as
unchanged and does **not** compile the branch)
- `pnpm check:changesets` — clean
- ~650 file-scoped tests including new negative-path suites for the
decision boundary, which previously had **zero** test coverage

`packages/engine/src/__tests__/plugin-runner.test.ts` fails 56/80 —
**verified pre-existing**, reproducing identically at base commit
`93a403af67` on `main`. Not in the merge gate.

### A mutation check that failed to fail

Worth recording, because it nearly shipped an untested security fix. The
first mutation check on the provisioning change reintroduced the `ceo`
hardcode and **all 17 tests still passed** — the tests asserted through
the policy path, which can no longer observe `isCallerPrivileged` at
all, precisely because `isPrivileged` is no longer forwarded there.
Org-chart cases that do exercise the function were added; the hardcode
now fails exactly 1 of 19, and restoring is green. A green mutation run
is only meaningful if the test can actually see the code under test.

## Known limitations (stated, not papered over)

- The bash containment floor is string-matching: a cost-raiser, not a
sandbox. Quoting, encoding, `$HOME`, symlinks, or an interpreter
one-liner can evade it. The durable protection is the decision route
refusing agent-originated deciders — the filter is the belt, not the
braces.
- Approval expiry is lazy (evaluated at decide/complete/redeem), not
swept, so an expired pending row stays visible in lists until touched.
- The extension's require-approval path returns a pending message but
cannot suspend a pi session mid-turn; engine-side pause hooks cover
engine lanes only.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Security**
* Hardened approval and permission gating with server-side decider
attribution, self-approval blocking, ownership checks, replay/race
protection, and status/TTL enforcement.
* Added fail-closed behavior for sensitive/unclassified tools and
sandbox provisioning approvals.
* Blocked credential/approval access via bash containment; plugin
destructive task operations now require explicit permission.
* **New Features**
* `fn serve` now defaults to bearer-token auth, with `--no-auth` as the
explicit opt-out.
* **Bug Fixes**
* Improved task move-source attribution (`moveSource: "user"`) and
tightened dashboard archive/bypass confirmation and operator attribution
behavior.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-30 21:50:37 -07:00
gsxdsm
f49e487d91 feat(core): untraited-project lane opt-in — and main was red on the FNXC gate (#2949)
Two things, and the second is why the first does not ship alone.

## The opt-in

`resolveProjectColumnsForRoles` gains `untraitedProject:
"declared-columns"`. When **no** workflow in the project expresses
**any** lifecycle trait, every declared column id joins the answer.

This is the three-state rule at **project** scope — the last item on the
deferred list, recorded at three self-healing call sites (#2869, #2876).
A board that renames its lanes and declares no traits contributes
nothing today, so its cards are **absent from every role-keyed query**,
and the correct per-card fallback downstream never runs for them. A
fallback cannot rescue a card the query never returned.

**Not "no workflow declares this role."** A project that expresses
traits and has no review lane has *answered*; widening there would
invent lanes it deliberately lacks. Mutation-verified both directions —
widening unconditionally fails 1 of 12, making the option a no-op fails
1 of 12.

**Opt-in, not default**, because the safe direction differs per caller —
the finding in `project-union-versus-per-task-lanes.md`:

| caller | over-inclusion costs |
|---|---|
| sweep | nothing — the per-card check discards the extra rows |
| aggregator | an inflated number an operator reads (#2864, #2866) |
| action site | a card routed or notified under a vocabulary that is not
its own (#2852, #2891) |

Making it the default moves all three at once, in the one direction two
of them must not. Verified byte-identical without the option, so this
lands with **no caller changes** and each site adopts it on its own
reasoning.

## Main was red, and my own gate caught me first

I dated the new comments `2026-07-31` while today is `2026-07-30` —
**the exact defect `check-fnxc-future-dates` exists to prevent,
committed while writing the feature.** The gate I added yesterday failed
my own commit.

Correcting mine surfaced that the merged sentinel batch, #2947, and
three engine test files carried future-dated stamps too, so **the gate
was failing on `main` for everyone**, not just here.

All corrected to real dates rather than raising the ceiling. The stamps
were simply wrong, and a baseline bump would have recorded the error as
permitted — which is the failure mode that ratchet exists to prevent.

Core and engine `tsc` clean, `pnpm lint` clean, census `--strict` 0,
FNXC gate 0 (469 known, none added), gate green (161/487/13/71).

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 20:17:18 -07:00
gsxdsm
c3df0f641b executor: orphaned tasks were never resumed after a restart on a renamed board (#2947)
`resumeOrphaned` is the only path that recovers tasks after a crash or
restart. On a board with renamed columns it recovered **nothing**.

## A missed pair, not an unconverted read

```ts
const tasks = await this.listWipLaneTasks();          // resolved by role — already converted
const inProgress = tasks.filter(
  (t) => t.column === "in-progress" && …,             // literal — discards everything the read found
);
```

The read was already resolved. The filter directly beneath it
re-asserted the literal on the rows that read returned, so the sweep
found the orphans and threw them all away.

**This is the worse half of the class, and it hid well:**

- the read *looks* converted, so scanning for `listTasks({ column: "…"
})` finds nothing;
- the census scores only the comparison, so the backlog number moves the
**wrong way** as you convert;
- a **structural test already existed** pinning "the read asks for
resolved lanes" — `executor-resume-query-lanes.test.ts` — and it was
green the entire time the sweep was dead. A test asserting the read
exists says nothing about the filter beneath it.

The failure only surfaces after a crash, when an operator is already
investigating the crash and has every reason to blame that instead.

## The ratchet, generalised

#2944 ratcheted this class inside `self-healing.ts` after review found
one instance and a follow-up audit found five more. This generalises it
to every engine source: a function that resolves lanes **and** compares
a column id in the same body is a pair.

Excluded, deliberately:
- the **fallback arm** of a resolved ternary (`lanes ? lanes.has(c) : c
=== "done"`) — the correct shape;
- four files whose literals are deliberate, each with the reason
recorded: `ephemeral-worker-manager` (unresolvable-workflow default),
`triage` (the U11 orphan case), `scheduler` and `replan-target` (sync
listeners on the inert sync IR reader, already pinned by
`sync-workflow-ir-is-always-default.pg.test.ts`);
- `self-healing.ts`, because it has a **dedicated** ratchet that is
strictly more precise. Two ratchets allowlisting the same site is one
fact with two owners, free to drift — the exact failure mode this
program keeps hitting. One file, one ratchet.

It carries a positive control: a wrong source path would make every case
pass by scanning nothing.

**I swept the rest of the engine with it and executor.ts was the only
genuine hit** — everything else is documented-deliberate or blocked on
the inert sync reader.

## Revert results

Each measured by restoring the literal filter and re-running:

| | reverted → |
| --- | --- |
| behavioural case | fails — the renamed card is dropped and the sweep
returns before touching it |
| the ratchet | fails, naming the site: `resumeOrphaned:
executor.ts:5974` |

A non-vacuous companion (card in the review lane → not resumed) rules
out a filter that matches everything: a card in review has no session to
resume, and re-dispatching it would restart finished work.

**Measured:** `executor.ts` column guards 8 → 7; baseline re-recorded
downward.

## Verification

`pnpm test:gate` 161 + 487 + 13 + 71; executor prompt/soft-delete/resume
suites plus the new ratchet, 357 passed; `tsc` engine clean; `pnpm
lint`, `check:changesets`, census `--strict` and
`check-sql-column-literals` clean, each run explicitly.
2026-07-30 19:59:03 -07:00
gsxdsm
9366bc8382 fix(workflow): the review handoff killed the walk on a renamed review lane (#2900)
The sharpest lane defect left in the backlog, and the one I have been
deferring since the first sweep.

```ts
if (seam === "review-handoff") {
  const result = await primitives.transitionTask(primitiveCtx, context.task, {
    column: "in-review",   // ← post-U12 this is a rejected destination on a renamed board
```

Post-U12 `moveTask` **rejects** a destination the workflow does not
declare. So on any board with a renamed review lane, the handoff threw
`TransitionRejectionError` and **killed the workflow walk mid-run**. Not
a silent wrong answer for once — a hard failure in the middle of a task,
which is why it outranked everything else once it became reachable.

**Why it was deferred:** every fix threads a resolver out of
`executor.ts`, and #2820 was editing that file. It merged at 22:08, so
this was finally free of the conflict.

## The role travels, not the column

Seam handlers in `workflow-node-handlers.ts` are pure functions over an
IR node and a task — no store, no task id to resolve from — so a handler
can only ever name a literal. The runtime primitive in `executor.ts`
**does** hold the store, so the seam now asks for `columnRole: "review"`
and the primitive resolves it against the task's **own** selection.

One authority, deliberately. Answering one question with two reads is
what took #2843 five review rounds, and I would rather not relearn it
here.

Compatibility is preserved in both directions:

- `column` still wins when both are supplied — an explicit destination
is an explicit destination;
- an unresolvable role falls back to the legacy `in-review` rather than
failing the transition, which is exactly the behaviour every caller had
before.

## The test asserts the literal is *gone*, not merely accompanied

`column` takes precedence over `columnRole` downstream, so a diff that
added the role while leaving the literal would look converted and be
completely inert. That is the exact shape this program keeps finding — a
documented fallback in front of a literal that still decides everything
— so the assertion is:

```ts
expect(input.columnRole).toBe("review");
expect(input.column).toBeUndefined();   // ← the half that matters
```

**Revert proof, measured:** restore `column: "in-review"` in the seam
and it fails with `expected undefined to be 'review'`.

## Verification

- `pnpm test:gate` — 161 / 487 / 13 / 71 passed
- `pnpm lint` — clean
- `tsc --noEmit` (`@fusion/engine`) — clean
- new `review-handoff-lane.test.ts` plus the two neighbouring seam
suites — 41 passed

Carries the one-line SQL-baseline re-record (`team-analytics.ts: 6 → 3`)
that #2864 left behind, same as my other open branches — main is red on
it, and identical changes to that line merge without conflict.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 17:39:59 -07:00
gsxdsm
89aaf341d0 the unwired-seam audit: 9 defects the census cannot see, incl. a reviewed card that cannot merge (#2820)
**Nine operator-visible defects in a class the census cannot see, plus
the audit method that found them.**

The census scans for lifecycle-column **comparisons**. This PR is about
guards that have no literal to find: a helper takes an optional
*resolved* lane set, its own test passes it, the census entry is gone —
and the callers pass nothing. **A resolved seam nobody wired is
indistinguishable from no seam at all.**

## What was broken

| defect | operator sees |
| --- | --- |
| `getTaskMergeBlocker` unwired in `mergeTaskImpl` | `Cannot merge FN-1:
task is in 'checking', must be in 'in-review'` — **a reviewed card
cannot merge** |
| …and in the completion move | `Cannot move FN-1 to done: …` — **and
cannot complete** |
| `isParkedTaskColumn` unwired ×2 (`agent-heartbeat`) | a durable agent
keeps claiming a parked card; **Health Check renders it RUNNING** |
| `resolveLinkSyncColumnRoles` first-per-role | link hygiene skips a
**second hold lane** entirely |
| `executor` active-task predicate first-per-role | a card in a **second
wip lane reads as INACTIVE**; its prompt file becomes reclaimable |
| `isPlanningContinuationTaskDispatchable` partially threaded | a board
declaring `done` as *non-terminal* stalls its cards — **stalled by a
lane name** |
| `default-workflow-hooks:72`, `executor:2404` | resolved gate admits
the move, unresolved blocker refuses it |

## The recurring shape, which is sharper than "a caller forgot an
argument"

Four sites resolve the lane and then re-ask with the literal, **a few
lines apart in the same function**:

- `task-artifacts-ops` resolves `completeColumn`, then asks the blocker
with the literal.
- `default-workflow-hooks:72` gates on `lifecycleColumns?.review`, then
the literal.
- `executor:2404` compares `resolveResumeLanes(…).review`, then the
literal.
- `resolvePlanningContinuationCandidate` applies the caller's terminal
set, then delegates without it.

**Grep for the helper, not the literal.** The literal is one function
away, correctly annotated as a fallback — which is exactly why the
census is blind to all of it.

## The arity trap, named and measured (six occurrences, one caught by
review here)

`resolveLifecycleColumns` answers *"which column is **the** hold
lane?"*. A `.includes()`/`.has()` test asks *"is this **any** hold
lane?"*. Nothing distinguishes them — same types, no literal.

**A default-vs-renamed differential cannot catch it**, because the
default board declares one column per role and therefore cannot express
the failing shape. It needs a *structurally* different fixture. That is
a sharper rule than "test both vocabularies", and it would have caught
all six.

Scanned: 12 candidate sites. **4 fixed · 3 blocked (2 on the inert sync
IR reader; `triage:833` also query-shaped) · 1 needs a hook-contract
change · 3 not defects (a returned tuple; an ordering-sensitive
precedence list) · 1 false positive of my own scan.**

A sweep over all twelve would have broken the ordering-sensitive pair,
delivered nothing at the sync-blocked ones, and "fixed" a site that was
already correct.

## Two traps in fixing this class — I hit both here

1. **The legacy id is a FALLBACK, not a member.** Pre-seeding
`"in-review"` admits a board that *declares* `in-review` as its WIP
column — a card mid-implementation merges prematurely. A real resolved
answer must **replace** the default. (Caught by review; it is the same
unscoped-legacy-acceptance the glasses plugin's review caught earlier,
which I had read and reintroduced.)
2. **Two guards, one assertion.** `toContain("must be in")` passed with
`mergeTaskImpl` reverted, because the *completion* guard caught the card
instead. The assertion now names the site (`Cannot merge` vs `Cannot
move … to done`) so the two fail independently.

## Corrections I made to my own work, recorded rather than quietly fixed

- My first PG test was **vacuous three ways**:
`saveWorkflowDefinition?.()`/`setTaskWorkflowSelection?.()` do not exist
(the `?.` swallowed both, so the task kept the builtin workflow),
`updateTask({column})` does not move a card, and a two-node IR made
every setup move illegal. Premise is now **asserted**, not assumed.
- My doc claimed the audit was complete. It enumerated **helpers**, not
every **caller** — `getTaskMergeBlocker` alone has 13 call sites.
Corrected in place, with the still-unwired ones listed by file and line
and a note to distrust any "audit complete" claim including mine.
- A severity correction to another worker's E2E:
`selectActionablePlanningContinuations` has **no production caller**, so
its stated consequence is latent, not live.

## Verification

- `pnpm test:gate` — 161 + 487 + 13 + 71
- `tsc` on core and engine; `pnpm lint`; `check:changesets`; census
`--strict` — all clean, each run explicitly
- Every fix revert-measured; each has a non-vacuous companion. The
two-hold-lane and repurposed-`in-review` cases exist because the default
board cannot express those shapes.

## Deliberately not done, with reasons in
`resolved-seams-nobody-wired.md`

`isTaskReadyForMerge` (dead in production — wiring it would be the
anti-pattern itself); `getTaskHardMergeBlocker` (3 of 4 callers are
query-gated sweeps); `getInReviewStallReason` (needs a **batch
prefetch**, not a per-task resolve — its callers decorate every task on
every list read; the in-review stall badge is wrong on renamed boards
until then); `default-workflow-hooks` planning/live-work sets (needs
`DefaultWorkflowMoveContext` to carry the IR — a shared contract
change).
2026-07-30 15:08:01 -07:00
gsxdsm
109204c590 fix: the query class — three sweeps that never ran on a renamed board (#2818)
Three sweeps that **never ran at all** on a renamed board, plus the
shared answer the rest of the class needs. Consolidated from three
handoff branches so the helper appears once. #2811 merged, so this is my
only open PR.

`#2800` measured this class and shipped evidence deliberately without
conversions: `listTasks({ column: "<literal>" })` filters in the store,
so on a renamed board the read returns an **empty array** and the sweep
it feeds does nothing. The census scores the comparison *inside* the
loop, never the query above it.

## What was broken

| file | census count | what actually happened on a renamed board |
|---|---|---|
| `backlog-pressure-reporter.ts` | **0** | both reads empty, ratio
computed as 0/0 — **the alert never fired**, on a board that may be
under exactly the pressure it reports |
| `stale-task-reporter.ts` | **0** | both reads empty — **no stale-task
signal ever raised**, where work is most likely sitting unnoticed |
| `restart-recovery-coordinator.ts` | flagged | sweep never ran — **an
engine restart left interrupted tasks stuck with no requeue** |

Two of the three have a census count of **zero**. They contain no
lifecycle comparison at all, so they have never appeared in the backlog,
in a per-file list, or in any "N → 0" claim — and were completely inert.
**A file at zero is not evidence of anything.**

## The shared answer, and what it is not

Every existing resolver answers a **per-task** question. A query has no
task in hand, so it needs the project-level one: every column any
workflow declares for a role, unioned with the legacy ids so a board
mid-rename still finds rows under the old ones. The set is never empty,
so a caller cannot accidentally query nothing.

The header states what it is **not**: answering a per-card question from
the union would mark a card as review because some *other* workflow
calls its column review — the flat-set mistake this program has made
four times.

## The finding that generalises: the query is rarely the whole defect

`stale-task-reporter` **still reported zero after the query was fixed**
— `getTaskAgeStalenessSignal` defaults to the legacy pair, so a card the
query now returned was refused inside the signal. Converting only the
query would have looked like a fix and changed nothing.

That is a caveat on #2800's approach, offered as refinement rather than
correction: **asserting the query ARGUMENT is right when pinning a known
defect** (the outcome is 0 either way) **and insufficient when proving a
fix**, because the outcome is the only thing that distinguishes a real
conversion from a deeper one. All three conversions here assert
outcomes.

`restart-recovery` had three layers — query, a redundant re-assertion
(deleted; a test pins the `paused` guard it did contribute), and a move
destination that was **already** resolved but whose warning comment was
stale. A stale warning is its own hazard: it told the next reader a
defect existed where none did.

## Verification

- helper **8 passed** · three reporter/coordinator suites **29 passed**
- `pnpm test:gate` **161 / 13 / 487 / 71** · lint clean · `--strict`
exits 0 · four `tsc` targets clean
- each conversion revert-proven independently; the failing case is named
in each test header

## Two mistakes worth recording

**The helper's own test caught a bug in it.** My first draft wrapped the
definition loop in one `try`, and `parseWorkflowIr` **validates** rather
than parses — one malformed row would have returned legacy-only lanes
for *every* workflow, indistinguishable from the bug it exists to fix.
Now isolated per definition.

**I clobbered the core barrel** by taking `index.ts` wholesale from a
handoff branch, dropping two exports `main` had added since; three
packages stopped compiling. Taking a file from another branch takes its
whole contents, including what is now stale — for a barrel that is
nearly always wrong. Re-applied as a single edit on top of `main`.

## Not included

`self-healing.ts`'s 49 — actively owned and mid-conversion; an outside
refactor there produces conflicting halves of one sweep.
`project-engine.ts` (7) and `executor.ts` (2) need their own read of
what each sweep does with the rows, which these three are the argument
for.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 13:55:10 -07:00
gsxdsm
6fc98fd6c7 the third census-invisible class: 51 hardcoded moveTask destinations, measured — and duplicates never archived on a renamed board (#2808)
A third census-invisible class, measured — plus the two worst instances
fixed.

## The shape

```ts
if (task.column !== "in-review") { … return; }     // the census counts THIS
await this.store.moveTask(taskId, "in-progress");  // and cannot see THIS
```

The census is an AST scan for **comparisons**. A `moveTask` destination
is a **call argument**, so no backlog entry ever points at one.
Converting the guard alone is *worse than converting neither*: the
handler starts admitting work on a renamed board and then tries to move
the card into a lane that board may not declare.

This bit twice in one week — #2797 (`branch-worktree` requeued into a
lane that may not exist) and #2807 (a GitHub "changes requested" review
dropped, then a move to a hardcoded `in-progress`). Both times it was
found only because the guard *next to it* happened to be under
conversion. So I went looking.

## Measured

Across `core`/`engine`/`dashboard`/`cli`/`plugins`, excluding
`__tests__`/`*.test.*` and comment lines:

| | count |
| --- | ---: |
| hardcoded `moveTask` destinations in production | **51** |
| …passing `recoveryRehome: true` — **deliberate**, not defects | 22 |
| …plain, rejected on a board that does not declare the target | **29**
|

**The 22 must not be "fixed".** `moves.ts` exempts them on purpose
(#1411): a card stranded in an undeclared column has to stay rescuable
to a legacy safe-landing column, or it can never be recovered at all. A
sweep that converts them deletes the rescue path. That distinction is
the reason this is 29 and not 51, and it is why I measured before
writing.

## Why this got sharper recently

The `workflowHasColumn(workflowIr, toColumn)` rejection used to sit
inside a block gated on `isWorkflowColumnsCompatibilityFlagEnabled` — a
settings key **nothing in production writes** — so it never executed and
the legacy `VALID_TRANSITIONS` table decided instead. U12 hoisted it out
of that dead branch and it is now live, proven on a real store by
`live-move-path-undeclared-target.test.ts`:

```
moveTask(card in "todo" -> "triage")  now REJECTS: /Unknown column for this workflow/
```

That changed the failure mode of all 29 from *"silently lands the card
in an undeclared column"* to *"throws"*.

**29 is not a crash count.** Whether a throw surfaces or disappears
depends on whether the caller catches, which is per-site and I did
**not** measure it — the doc says so explicitly rather than letting the
number imply severity it hasn't earned.

## Fixed here: 9 of the 29

`duplicate-intake` and `duplicate-guard` both archive a duplicate. On a
renamed archive lane the move is rejected, so **the duplicate is never
archived and keeps sitting on the operator's board as live work** — and
in `duplicate-guard` the row has already been stamped
`deterministicDuplicateOf`, so it is *marked* a duplicate while
occupying an active lane. Half-applied, which is the same trap as
#2797's branch clear.

Both now resolve the `archived`-trait column from the task's own
workflow through one shared helper, unioned with the legacy id.

**`cli/commands/task-lifecycle`** — `finalizePullRequestMerge` and
`finalizeNoOpMergeTask` both move the card to a hardcoded `"done"`, and
both run `updateTask({ status: null, mergeRetries: 0 })` *first*. On a
rejection the merge has already landed and the bookkeeping is already
cleared while the card never reaches its complete lane: the operator
sees a merged branch, a card still sitting in review, and a reset retry
counter. Same half-applied shape as #2797's branch clear. Both now route
through one resolver so they cannot drift.

**`contamination` / `foreign-only-contamination` (×2) /
`restart-recovery-coordinator`** — four recovery requeues to a hardcoded
`"todo"`, none of them a `recoveryRehome` escape. On a board without
that column the move is rejected and **the recovery never completes** —
the card stays contaminated or stranded, which is precisely the state
these paths exist to clear.

**Consolidation.** `resolveReboundTargetForTask` and
`resolveArchiveTargetForTask` now live beside
`resolveTaskLifecycleColumns` in `workflow-lifecycle-traits`, already
the store-dependent resolution seam. My first pass put the archive
helper inside `duplicate-intake` and had `duplicate-guard` import it
from there — wrong home, and it would have grown a copy per caller as
more sites converted. Seven call sites now share two definitions.

**Plain (non-`recoveryRehome`) destinations: 29 → 21.**

**Coverage on the CLI pair is scoped, and I'd rather say so than imply
more:** the test covers the *resolver*, not the two call sites. Both
enclosing functions are private and reachable only through
`processPullRequest`, which needs a live GitHub surface — exporting them
purely to test wiring is a worse trade than stating what is covered.
Three cases: renamed lane resolves, no-workflow falls back to the legacy
id (which also pins that a default board is byte-identical), and a
throwing lookup falls back.

## Revert result (measured)

| conversion | reverted → |
| --- | --- |
| duplicate archive destination | new case fails — `moveTask` called
with `"archived"` on a board whose archive lane is `boxed` |
| CLI complete-lane resolver | replacing the body with a bare `return
"done"` fails the renamed case |
| both move-target resolvers | replacing either body with a bare return
of its legacy id fails 5 cases across the resolver suite and
`duplicate-guard` |

Each resolver has a **non-vacuous companion** asserting it does *not*
return the legacy id on a renamed board — without it, a resolver
returning any string would pass. The fallback cases are load-bearing
rather than padding: `resolveWorkflowIrForTask` degrades to the built-in
IR rather than throwing, and the built-in rebound/archive lanes *are*
`todo`/`archived`, so those cases also pin that a default board is
byte-identical.

The pre-existing case asserting the legacy `"archived"` passes both
ways, which is exactly why it could not detect this and why the new one
supplies a workflow.

## Ownership note

`packages/core` was `batch-core`'s territory and `packages/cli` was
`batch-cli-plugins`'. Both batches have landed, and this is
newly-discovered work in the class documented here rather than leftover
conversion backlog. Four sites, two shared helpers — happy for either
half to move if those owners would rather carry it.

## Verification

- `pnpm test:gate` — 161 + 487 + 13 + 71, green
- `duplicate-guard` + `duplicate-intake` — 40 passed
- `tsc` on core and engine — clean
- `pnpm lint`, `check:changesets`, census `--strict` — all clean (run
explicitly; a clean `pnpm lint` alone is not evidence the CI Lint check
passes)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

- **Bug Fixes**
- Duplicate tasks are now archived to each workflow’s configured archive
lane.
- Completed tasks are moved to the workflow-specific completion lane,
with a safe fallback for older workflows.
- Recovery and requeue actions now use each workflow’s configured
rebound lane instead of assuming a fixed destination.

- **Documentation**
- Added guidance on avoiding failures caused by hardcoded workflow
destinations and incomplete lifecycle conversions.

- **Tests**
- Added coverage for renamed workflow lanes, fallback behavior,
duplicate archiving, and recovery destinations.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 12:25:00 -07:00
gsxdsm
e84e9d7f60 fix: the caller audit — five unwired parameters, five defects in their callers (#2803)
Seven fixes that were sitting on separate handoff branches with no owner
while `main` moved. Consolidated, rebased onto current `main`, and
verified **together** rather than only per-branch. The individual
branches remain if a subset is preferred.

This is the same consolidation that got `batch-core` and #2787 adopted.
**Close it if it breaks queue policy** — the branch keeps the work safe
either way.

## Where these came from

#2787's review found an optional parameter whose production caller never
passed it. That is a class, so I ran it against everything I had landed
and found five more. **All five turned out to have their real defect in
the CALLER, not the parameter** — in four of them the parameter was
unreachable:

| unwired parameter | what was actually wrong |
|---|---|
| `blocker-fanout.escalationColumns` | the hold default made the count
zero — **no bottleneck warning was emitted at all** |
| analytics `columnFlagsByName` | routes never built a map — **0
in-progress / 0 in-review beside correct cost totals** |
| `isLegacyAutoMergeStampCandidate` | the read **queried a column a
renamed board does not have**, so the backfill iterated nothing |
| `rankAssignedTasksForWakeDelta` | `getTasksByAssignedAgent`'s
`excludeArchived` used the literal — **archived cards returned as open
work** |
| `duplicate-intake.columnFlagsByColumnId` | intake could **archive or
soft-delete a newly created task** as a duplicate of finished work |

The heuristic worth keeping: **an optional parameter no production
caller fills is a marker pointing at an unexamined caller.** The census
cannot see any of these five — every gate is a `Set`/array literal or a
query filter, i.e. a definition rather than a comparison.

## Also included

- **`executor.ts`** — the stale-spec guard did the exact thing its own
comment forbids: on a renamed board it ran on a LIVE task and pulled it
out of execution into replan. `activeMergeStatuses` protected merging
cards *by accident*, which is why the symptom looked arbitrary.
- **`register-project-routes.ts`** — project health reported **0 active
tasks**; its list also still contained `triage`, dead since U11.
- **`dashboard/app/utils/taskTiming.ts`** — a **second copy** of
`getTotalAgentActiveMs`. Core's was converted; the card chip imports
this one, so the census counted the site as done while the rendered
number stayed keyed on `"in-progress"`.

## Verification

Verified as a set: `pnpm test:gate` **161 / 13 / 487 / 71** · core
suites **15 passed** · engine **7** · dashboard **12** · four `tsc`
targets clean · lint clean · census `--strict` exits 0.

Each fix is revert-proven individually; the specific case that fails is
named in each test header.

## Two honesty notes

**Three guards here are structural, not behavioural, and say so in their
headers.** `sanitizeAgentTaskLinks` is a closure inside
`createApiRoutes`; the analytics aggregators need a live
`AsyncDataLayer`; the stale-spec guard sits deep inside `execute()`.
Each ratchet fails on revert — verified — but none is an end-to-end
proof, and the headers state which half they cover.

**One of my behavioural test sets would have lied.** The intake-dedup
cases drive `findSameAgentDuplicates` directly; I removed the wiring to
measure the revert and **they stayed green**, because they pin the
predicate and not the caller. That is the exact illusion this audit was
chasing, reproduced in my own file. The forward now has its own
structural check.

## Deliberately not included

`worktree-pool.ts:1205` — it **fails safe** (a missed match protects a
branch from cleanup rather than deleting it) and sits in the merger's
branch-reaping path where the opposite error destroys work. That
deserves its owner's judgement, not a drive-by conversion.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 12:02:53 -07:00
gsxdsm
a8cfce8fbd executor: stale merge evidence re-entering execution, and a live checkout that read as unowned (12 → 8) (#2805)
Two executor conversions with real operator consequences, one census
false positive, and three sites deliberately left with their reasons
recorded.

## Census

| file | main | here |
| --- | ---: | ---: |
| `packages/engine/src/executor.ts` | 12 | **8** |

Of the 4: three genuine conversions, one reclassification.

## What was broken

**`resetMergeStateIfNeeded` — cards re-entered execution carrying stale
merge evidence.** Merge state is cleared when a card *leaves* a lane
where a merge could have been recorded. Keyed on `in-review`/`done`, a
renamed board matched neither, so a card bouncing back into execution
kept `mergeDetails` — including a **commit sha from its previous pass**
— into its next run. `review` is not a trait, so this resolves through
the same five flags (`complete`, `mergeOrchestration`, `mergeBlocker`,
`humanReview`) the dependency gates in this file already use; two gates
answering "is this a merge-bearing lane?" differently would be a split
brain.

**The worktree-owner scan — a live checkout read as unowned.**
`findActiveWorktreeOwner` asks "is anyone else working in this
checkout?". Its in-memory `activeWorktrees` leg is
vocabulary-independent, but the **durable** leg — the one that answers
after an engine restart, when the in-memory map is empty — filtered with
`t.column !== "in-progress"`. On a renamed board that matched nobody, so
the worktree read as free and a second task could be handed a checkout
another task is live in. Post-restart is exactly when this function
matters.

Not the query-filter class: that `listTasks` call passes no `column`, so
the predicate is the only lane gate on the path.

## A third census false positive in this package

Line 16094's `to` is a **review-addressing record status** — the method
signature is `to: "queued" | "in-progress" | "addressed" | "failed"`,
and the next two lines test it against `"addressed"` and `"failed"`,
which are not columns at all. Marked `DELIBERATE-LITERAL`.

That is the third in `packages/engine` after the two `cli-agent`
`CliMachineState` ones (#2797). The backlog total includes non-columns;
a sweep that "converts" them turns a status machine into a workflow
role.

## Revert results (measured, each run independently)

| conversion | reverted → |
| --- | --- |
| worktree-owner wip predicate | RENAMED case fails — checkout reads as
**free** while another task is live in it |
| `resetMergeStateIfNeeded` lanes | RENAMED case fails — card keeps
`commitSha: "abc123"` from its previous pass |

Both DEFAULT cases pass before and after, which is why both vocabularies
run. Each has a non-vacuous companion (holder sitting in the complete
lane; a return from the hold lane) so a predicate matching every column
would not pass.

**Both reach their seam directly through a cast.** The public routes are
`handleBranchConflict` (needs a real `BranchConflictError` plus a git
repo) and the `task:moved` listener (drags in the whole `execute()`
path); going through either would make these tests about a git fixture
rather than about the lane predicate. The alternative was the status quo
— all 91 `executor-worktree*.test.ts` cases seed `column:
"in-progress"`, so they assert the legacy fallback and pass either way.
I shipped the conversions in one commit *stating* they were unproven,
then closed that gap in the next; the history shows both.

### Two fake defects found while writing those tests

Worth naming, because both are the documented green-for-the-wrong-reason
shape:

1. The first fake had no `updateTask`, so the cleanup **threw** rather
than asserting anything.
2. The second returned a new object without persisting — and
`cleanupMergeStateForReverification` **re-reads through `getTask`**. The
re-read handed back the stale row, so *both* vocabularies reported
"nothing changed" and it would have read as a passing negative test.

## Deliberately NOT converted, with reasons

- **The `task:moved` listener cluster** (`3521`/`3545`/`3596`/`3606`),
including the AGENTS Move-Task hard-cancel contract `userCanceled:
source === "user" && to === "todo"`. Its prologue is synchronous
(`userCanceledTaskIds.delete`, watchdog clear) and deferring it to a
microtask changes hard-cancel ordering. The sync IR reader is not an
option — it returns the DEFAULT workflow for every task in production. A
safe conversion needs lanes resolved on an earlier async boundary: new
machinery plus an ordering change, which is out of fleet scope and not a
guess worth making on a hard-cancel path.
- **`17081`** pairs `latestColumn === "in-progress"` with a
**hardcoded** `moveTask(taskId, "in-progress")` two lines above —
census-invisible, the same shape as the branch-worktree requeue bug in
#2797. They have to convert together, and the move needs the same
rejection guard.
- **`5903`** is the query-filter class: `listTasks({ column:
"in-progress" })` followed by a re-assertion of the same literal.
Converting it drops a count and changes nothing — see
`docs/solutions/architecture-patterns/self-healing-sweeps-are-blind-on-a-renamed-board.md`
(#2800).

## Verification

- `pnpm test:gate` — 161 + 487 + 13 + 71, green
- `npx tsc -p packages/engine/tsconfig.json --noEmit` — clean
- `pnpm lint` — clean
- `--strict` exits 0

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 12:02:41 -07:00
gsxdsm
90f6319b79 batch-engine tail: re-land the ASYNC half; the sync-resolved half was inert (engine −15) (#2785)
Tail of `batch-engine` (#2773). That PR merged as a squash while later
engine work was still in flight, so `self-healing.ts`, `executor.ts` and
`worktree-pool.ts` landed at their pre-conversion counts. This re-lands
**only the half that is real**, and the reason the other half is not
here is the substance of this PR.

## Census, per file (measured, `--strict` verified)

| file | main | here |
| --- | ---: | ---: |
| `engine/src/self-healing.ts` | 107 | 97 |
| `engine/src/executor.ts` | 15 | 12 |
| `engine/src/worktree-pool.ts` | 3 | 2 |
| `engine/src/ephemeral-worker-manager.ts` | 1 | 0 |
| `engine/src/agent-tools.ts` | 5 | **0** |
| `engine/src/gridlock-detector.ts` | 3 | **0** |
| `engine/src/triage.ts` | 4 | 1 |
| `engine/src/mission-execution-loop.ts` | 2 | **0** |
| **net** | | **−28** |

Baseline re-recorded; `--strict` tightened exactly these 4 entries and
no others.

## Finding: a whole class of conversions in this program is INERT, and
the census scores it as progress

`resolveTaskWorkflowIrSync` returns the **default** workflow IR for
every task in production. The sync selection reader behind it is a
PostgreSQL-cutover stub:

```ts
// packages/core/src/task-store/workflow-definitions.ts:505
export function getTaskWorkflowSelectionImpl(_store, _taskId) {
  return undefined;   // "Backend mode cannot synchronously read PostgreSQL"
}
```

So a guard written as
`resolveLifecycleColumns(store.resolveTaskWorkflowIrSync(id))?.hold`
resolves an IR, asks for a trait, and answers **from the default
workflow for every custom board** — silently. It reads as converted and
the census counts it as converted. `main` gained
`sync-workflow-ir-callsite-allowlist.test.ts` for exactly this after my
branch point; it is what caught me.

I had built three sync resolvers on that reader — `resolveMoveLanesSync`
(self-healing, executor) and a widened `resolveTaskParkedColumnsSync`
(scheduler) — reasoning that a *synchronous* `task:moved` listener needs
a *synchronous* reader. That reasoning was sound about the shape and
never checked whether the reader reads anything.

**Dropped from this PR, deliberately, and NOT re-landed anywhere:**

- `scheduler.ts` 12 → 1 (the widening; the pre-existing narrow helper on
main is untouched)
- the executor `task:moved` handler, incl. the Move-Task hard-cancel
lane comparison
- self-healing's `task:moved` fan-out,
`classifyPausedAbortWorkflowRecovery`, `reconcileInReviewBranchRebind`,
`recoverWedgedActiveMerge`, `recoverPausedAbortFailures`, and 12
single-row lane conversions

Those sites are back to their literals. The allow-list's own guidance is
the standard I applied:

> An unconverted `=== "todo"` is strictly better, because it is at least
honest about being a literal.

I did not add my call sites to the allow-list. Six entries would have
turned the gate green in two minutes and buried the defect; the list's
contract requires proving the async resolver is genuinely unreachable,
and for a fire-and-forget listener it is not — the listener can `void`
an async lane resolution the same way `NotificationService` already
does. That is the correct fix and it is a behaviour-shaped change, so it
is out of scope here.

**Fleet-wide consequence:** any conversion routed through
`resolveTaskWorkflowIrSync` is fake progress, and the census cannot see
the difference. `pnpm test:gate` can: the allow-list test is the
detector. Its passing here (161/161) is this PR's evidence that nothing
inert survived the split.

## What IS in this PR — all async-resolved

1. **`self-healing.clearStaleBlockedBy`** — lanes resolved per
**REFERENCED** task, not per iterated task. A blocker's own workflow
decides whether it is still blocking.
2. **`executor` dependency satisfaction** — resolved per **DEPENDENCY**
via `columnsWithFlag`. Preserves the load-bearing asymmetry that a
dependency in *review* already satisfies a dependent; a bulk sweep
flattens that to complete-only and deadlocks the board.
3. **`agent-tools` — the agent task tools listed FINISHED cards as
active.** `fn_task_list` says it lists "tasks that aren't done or
archived"; `fn_task_search` offers `includeDone: false`. Both filtered
on `task.column !== "done"`, so a renamed complete lane returned
finished cards as outstanding work **to an agent**, which then reasons
and acts on them. `includeArchived` was always enforced by the QUERY and
survived a rename; `"done"` was only ever a TS predicate, which is why
exactly that half broke.

Plus the two **dedup** guards in the same file. The cross-parent
diagnostic filter kept a *shipped* card as a candidate on a renamed
board, so the guard adopted it as canonical and returned `wasDuplicate:
true` — absorbing new diagnostic work into a task nobody is working on
(the eval-followup defect shape again). The defined-feature bootstrap
preflight is **not** the query-filter class: its query passes
`includeArchived: true`, so the TS predicate is the *only* archived
guard there; on a renamed archive lane the archived sibling became the
bootstrap canonical and `claimDefinedFeatureTask` then rejects the
non-live row, so a valid first task fails to be created at all.

Both dedup invariants **already had tests** — asserted against the
legacy ids only, so both passed for the very comparison being replaced.
Extended in place into vocabulary differentials rather than added as
parallel files. Two helpers rather than one parameterised one: "is this
finished?" and "is this archived?" are different questions, and merging
them would make the archived-only guard also reject completed rows.

The list/search half re-landed **with the test it originally shipped
without.** No suite exercised either tool, so the original commit's
"304/304 green" said nothing about the change — the optional-flags
failure mode exactly. Both call sites are covered; converting two copies
and testing one is the Surface Enumeration failure this program has
already hit twice.

4. **`gridlock-detector` — FALSE dependency alarms.** The gate compared
each blocker against `done`/`in-review`/`archived`; on a renamed board
all three are true for a *finished* blocker, so no dependency ever
counted as met and the detector reported dependency gridlock for tasks
that are not blocked — `notifyGridlock` then pages the operator.
Resolved per dependency using the **same five flags** as the executor's
gate (`complete`, `archived`, `mergeOrchestration`, `mergeBlocker`,
`humanReview`) — `review` is not a trait, and two gates answering "is
this dependency satisfied?" differently is a split brain. Every
pre-existing case in that file omits a workflow, so none could detect
the change; added the renamed case plus a non-vacuous companion.

5. **`triage` — its OWN copies of the same two tools.**
`createTriageTools` carries a `fn_task_list` and `fn_task_search`
byte-identical in intent to the agent-tools pair, plus a third site
filtering duplicate candidates. Same defect on all three. Reused the
(now exported) agent-tools helper rather than adding a third copy —
deliberately stronger than the two-parallel-tests reading of Surface
Enumeration, since the copies now share one implementation and cannot
drift. **Not claiming call-site coverage:** `createTriageTools` is
private and not drivable without standing up a TriageAgent; the helper
is revert-proofed, those two call sites are covered only through it.

6. **`mission-execution-loop` — a finished fix task read as LIVE,
stalling remediation.** The comment above that line states the rule it
implements: *only an open task makes duplicate triage safe to suppress.*
On a renamed board the rule inverts — a finished fix task is not
`done`/`archived`, so it reads as live, remediation for a fresh
validation failure is suppressed indefinitely, and the mission stalls
with no error surfaced.

**Not revert-proven, and I am not claiming it is.** No test reaches the
`hasLiveFixTask` branch, and the only case that mints a fix feature is
git-gated and heavyweight; building that fixture is larger than the
conversion. The change strictly *widens* the finished set (resolved
roles ∪ the two legacy ids), so default boards are byte-identical — that
is the argument for shipping it unproven, not a substitute for coverage.

7. **Four census-invisible membership guards**, each inverted on a
renamed board — `worktree-pool` (merger-managed branch reclaim could
delete a branch out from under an in-flight merge), `agent-assignment`
(assignment load counted nothing), `ephemeral-worker-manager`
(`isAgentIdle` inverted on both sides), and the dead constants their
conversion orphaned. These are `SET.has(task.column)` shapes the census
does not count, so the −15 understates them.

## Revert results (measured, each run)

| conversion | reverted → |
| --- | --- |
| `clearStaleBlockedBy` per-referenced lanes | renamed-vocabulary case
fails; stale `blockedBy` never clears |
| executor dependency satisfaction | dependent never unblocks on a
renamed review lane |
| `worktree-pool` merger-managed set | reclaim proceeds against an
in-flight merge |
| `ephemeral-worker-manager.isAgentIdle` | idle agent reads busy on a
renamed board |
| `fn_task_list` terminal filter | RENAMED case fails — shipped card
listed as active |
| `fn_task_search` terminal filter | RENAMED case fails — same,
independently |
| cross-parent diagnostic dedup | RENAMED case fails — `wasDuplicate:
true`, new work absorbed |
| bootstrap preflight archived guard | RENAMED case fails — `validate`
called with the archived sibling |
| gridlock dependency gate | RENAMED case fails — false gridlock raised
for an unblocked task |

`agent-assignment`'s widened `taskStore` type is compile-time; its
revert is a tsc failure, not a test failure — stated rather than claimed
as coverage.

## Verification

- `pnpm test:gate` — 161 + 487 + 13 + 71, all green (161 includes
`sync-workflow-ir-callsite-allowlist`)
- `npx tsc -p packages/engine/tsconfig.json --noEmit` — clean
- `pnpm lint` — clean

One commit is a pure import restore: `columnsWithFlag` arrived in a
sibling commit that built on the inert resolver and was left behind. The
engine tsconfig excludes `src/__tests__/**`, so the gate was green while
tsc was not — worth knowing that on this package a green gate is not a
green build.


## Verified NOT a gap — measured, so the next worker does not re-open
them

- **`restart-recovery-coordinator` (5 counted).** Four already take an
optional `reviewColumns` set and the counted literals are the documented
**fallback** arm, which must stay for the same reason `columnRoles.ts`
keeps its id fallback. The sole production caller
(`self-healing.ts:12151-12154`) already passes the resolved set. The
fifth is documented at the site as a re-assertion behind a `listTasks({
column: "in-progress" })` query filter. Nothing to convert.
- **`notification/notification-service` (5 counted).** Already
documented in-file as deliberately counted with no exemption marker: the
wedge-episode site needs per-task serialisation of wedge handling (a
delivery-semantics change to operator notifications), and
`isManualMergeHold` needs a pre-resolved `LifecycleColumns` threaded
through `handleTaskUpdated`, which would pay resolution on every task
update. Both are behaviour/placement judgements, not conversions.
- **`planner-overseer` (3 counted).** `resolveWatchedStage`'s two
literals are fed by `pollPlannerOverseer`, which calls `listTasks({
column: "in-progress" })` and `{ column: "in-review" }` — hardcoded
**query** filters. On a renamed board those queries return no rows, so
the predicate never sees a renamed column. Converting it alone would
drop 3 from the census and change nothing an operator can observe. The
real fix is at the query layer; that is the tracked query-filter-bounded
class, not this PR.
- **`triage:695`** reads `resolvePlannerLanes` → the allow-listed sync
IR reader. Left as an honest literal per the rule above.

**Still open in `packages/engine`, deliberately not in this PR:**
`self-healing.ts` (97, of which ~31 are the query-filter-bounded class
and the rest need per-site classification in a 13k-line file),
`scheduler.ts` (12, blocked on the sync reader above), `executor.ts`
(12), and a tail of ~13 more copies of the "is this task finished?"
question across eight small files (`agent-reflection`,
`auto-merge-finalization`, `merger-scope-auto-widen`,
`backlog-pressure-reporter`, `merger-orphan-rehome`,
`merger-integration-worktree`, `plugin-runner`, `cli-agent/*`). That
tail is a clean follow-up: one question, eight call sites, and the
exported `resolveTerminalColumnsForTasks` helper already exists for it.

That is the same discipline as the sync-resolver finding: a census
number that drops without a behaviour change is not progress, and four
of these files would have handed over exactly that.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 10:47:44 -07:00
gsxdsm
6bdde6f246 fix: five lifecycle gates the census cannot see — incl. live ephemeral workers reaped and duplicate follow-up cards (#2787)
Five lifecycle-column fixes the census **structurally cannot see**. Each
gate is a `Set` or array literal — a *definition*, not a comparison — so
no backlog entry ever pointed at any of these files. Found by grepping
for lane-shaped list literals after the same shape surfaced in
`duplicate-intake` and `blocker-fanout` (both merged via #2780), then
confirmed by reading each USE site.

**On opening this:** I offered twice to fold these into a PR and kept
them on handoff refs to respect one-open-PR-per-worker. They have now
sat unadopted across several cycles while `main` moved, and two of them
destroy or duplicate work. Opening is the reversible call — **close it
if it breaks queue policy** and I will keep them on the branch.

## What is in it

| commit | defect on a renamed board | severity |
|---|---|---|
| `beb107a7bc` | assignment load-balancing **defeated** —
`assignmentLoad` stays empty, every candidate reads as load 0, the sort
falls through to its stable `createdAt` tiebreak, so **one agent wins
every assignment** while the rest idle | distribution |
| `cf4b59e1cb` | the zombie sweep **deletes LIVE ephemeral workers** |
**destroys work** |
| `5fe004ae64` | eval follow-up dedup sees **zero open tasks**, so every
run re-files follow-ups it already filed | **duplicate cards** |
| `a1021de8b2` | agents keep a **"working on" indicator for finished
cards** | stale UI |
| `86680d1220` | the **Files tab never loads** — the fetch never fires |
silent empty |

### The one that destroys work

`shouldDeleteOnSweep` tested a hard-coded terminal `Set`, then fell
through to `return task.column !== "in-progress"`. On a renamed board
**both halves miss, and they compound in the worst order**: the terminal
test fails, control reaches the fallthrough, and `"building" !==
"in-progress"` is `true`. An ephemeral worker **actively executing a
task** is classified as a zombie and deleted. Nothing logs.

Its fallback is **deliberately asymmetric**, and the comment says why:
an unresolvable workflow keeps the legacy literals rather than guessing.
Failing to reap a dead worker costs a slot; reaping a live one destroys
work in flight. Those are not symmetric, so uncertainty fails toward
keeping the worker.

## Verification

Verified **as a set**, not only per-branch:

- `pnpm test:gate` — **161 / 13 / 487 / 71**
- engine suites (assignment, ephemeral, eval-followups) — **44 passed**
- dashboard suites (agent-task-link, useSessionFiles) — **16 passed**
- `tsc` engine + dashboard server + dashboard app — clean
- `pnpm lint` clean · census `--strict` exits 0

**Revert-proven individually.** Restoring each literal fails its own
case: the renamed-wip zombie case, the renamed-wip assignment case, the
renamed-lane dedup case, the sanitizer ratchet, and both
`useSessionFiles` role cases.

## Two honesty notes, flagged rather than buried

**`a1021de8b2`'s guard is STRUCTURAL, not behavioural.**
`sanitizeAgentTaskLinks` is a closure inside `createApiRoutes`,
reachable only by standing up the full express app. The ratchet asserts
the source — resolver threaded per task, bare literal call gone, cache
shared, fallback retained — and **fails on revert**, verified. It is not
a substitute for a behavioural test; whoever owns the dashboard server
should add one if that seam grows.

**`useSessionFiles`'s negative case passed in isolation and failed in
the suite.** Hooks are not unmounted between cases there, so a prior
case's in-flight fetch landed inside it. That is the classic shape of a
test that gets "fixed" by reordering; it now asserts a **delta** against
the pre-render call count, which is independent of what leaks in.

## Deliberately NOT included

`worktree-pool.ts:1205` — the sixth site from the same sweep. It **fails
safe**: a missed match means the skip does not fire, so the branch is
added to `activeBranches` and *protected* from cleanup. The cost is
stale branches accumulating, not deletion. It also sits in the merger's
branch-reaping path, where the opposite error destroys work, so it
deserves its owner's judgement rather than a drive-by conversion.
Flagged, not guessed.

Also still open and unclaimed: roughly 69 untriaged literal-list sites
across engine/dashboard/cli. The grep is one line and the file list is
on #2775 — with the measured caveat that about half are false positives
on shape alone (`LEGACY_*` names, seeds unioned with resolved values,
and `roles: ["triage"]`, which is an `AgentCapability`, not the deleted
column). Only the use site settles it.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 10:35:31 -07:00
gsxdsm
c643d62e85 fix(executor): wipDeclared must ask ALL six lifecycle roles, not two (#2777)
## What this fixes

`resolveResumeLanes` returns `wipDeclared`, which gates whether
`routeGraphFailureToExecutionResume` may route a graph failure back into
execution resume. Getting it wrong terminalizes tasks on boards that
should resume.

Two prior versions were wrong, both caught in review rather than by me
at write time:

**1. Two-state (`lifecycle?.wip !== undefined`)** — greptile P1 on
#2760. A v1-upgraded board terminalizes: `synthesizeDefaultColumns`
emits `{ id, name: id, traits: [] }`, so *every* role resolves
`undefined` even though those columns literally are the legacy lanes.
Verified by parsing a real v1 IR.

**2. Proxying "synthesized" as "hold and review are both undefined"** —
my own fix for (1), and also wrong. I caught this against #2765 rather
than shipping it. A **v2** board that declares only `intake` +
`complete` has hold and review undefined too, so it would be misread as
synthesized and treated as declaring wip when it deliberately does not.

The failure mode both versions share: reading a *sample* of the roles
and treating the answer as a verdict about the whole IR. #2765 says it
directly — an empty result has two meanings, and you cannot tell them
apart from a subset.

## The rule

```ts
wipDeclared: lifecycle?.wip !== undefined || !declaresAnyLifecycleRole(lifecycle),
```

Three states, asking all six roles:

- **wip declared** → true, the board says so.
- **some role declared but not wip** → false. A v2 board that omits wip
means it; do not resume into a lane it did not define.
- **no role declared at all** → true. That is the
synthesized/v1-upgraded shape, whose columns *are* the legacy lanes; the
pre-existing behaviour is correct there and must not regress.

`declaresAnyLifecycleRole` iterates `Object.values(lifecycle)` rather
than naming roles, so a seventh role added later is included
automatically instead of silently falling into the wrong branch.

## Evidence

- `executor-resume-lanes-resolved.test.ts`: **7 passed**, +23 lines
covering the v1-synthesized board and the declares-some-but-not-wip
board.
- **Mutation:** restoring the naive two-state rule → **1 failed / 6
passed**. The added coverage is load-bearing and pins exactly the
regression greptile caught.
- Gate **732 green** · `pnpm lint` clean · engine `tsc --noEmit` **0
errors**.
- Rebased on current main.

## Scope

`executor.ts` (+22) and its test (+23). One predicate; no other behavior
touched.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-07-30 09:28:41 -07:00
gsxdsm
7119432c79 fix(engine): the stranded-completed recovery never resolved on a renamed board — and its suite fed the broken reader the right answer (#2764)
## The "recovery of last resort" never resolved anything

`recoverCompletedTask` carries this note in its own source:

> *This is the recovery of last resort — a literal here means the last
resort does not exist off the default lineage.*

It was resolving through `resolvePlannerLanes`, which reads
`resolveTaskWorkflowIrSync` — whose selection reader returns `undefined`
**unconditionally** in PostgreSQL mode, the shipped backend. So it
resolved the **default** workflow for every card,
`promotedFromPlannerColumn` was `false` on every renamed board, and the
recovery never fired.

That is precisely the stranding it exists to fix — completed work
sitting in a planning lane with nothing left to rescue it — **with the
conversion in place and the census counting it as done.**

The call site is inside an async method that has already awaited store
reads, so the fix is an `await`, not a restructure.
`resolvePlannerLanesForTaskAsync` is the async twin: identical logic,
identical fallbacks, one `await`. Answers are unchanged on the default
lineage and correct everywhere else.

## The existing suite could not see any of it — the more important half

`executor-planner-lanes-resolved.test.ts` injected **only**
`resolveTaskWorkflowIrSync`:

```ts
(store as { resolveTaskWorkflowIrSync: ... }).resolveTaskWorkflowIrSync = () => ir;
```

It fed the broken reader **the right answer**. Every case proved the
promotion *logic* while being structurally blind to whether production
resolves at all — and it was green the entire time. A suite that cannot
fail for the reason the code is broken is the same defect as the code,
one level up.

The harness now feeds the sync reader the **default lineage** (what it
actually returns) and the async readers the task's real workflow.

| | reverting the call site to sync |
|---|---|
| before this PR | **0 failed** — suite blind |
| after | **5 failed** / 13 passed |

Two cases opt back in via `syncResolvesIr`, and only those two: they
cover `isPlannerColumnFor` and `isBackwardMoveOutOfPlanning`, which are
still synchronous, so there the sync reader genuinely *is* the input
path and feeding it the IR tests the classifier rather than the reader.

## Not converted, deliberately

**Those two classifiers.** They sit in an else-if chain whose next arm
is `from === "in-progress"`, so deferring the decision into an async
body changes which arm runs. That branch's own comment records a
previous half-conversion there:

> *a half-conversion turned a missed rescue into active damage. Third
time this program has produced that shape — gates converted,
destinations left literal.*

That needs the chain enumerated first, not a fast restructure at the end
of a sweep. Their production inertness is held by the
`resolveTaskWorkflowIrSync` call-site allow-list in #2759, so they
cannot be forgotten.

## Verification

- new suite **4 passed** · strengthened suite **14 passed** (18
together)
- `pnpm test:gate` — **158 / 10 / 487 / 71** · `pnpm lint` clean ·
engine `tsc --noEmit` **0 errors** · `--strict` exits 0

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 08:38:43 -07:00
gsxdsm
d86c1f9d29 batch-engine: packages/engine lifecycle-column conversions (capacity worker's mega-batch) (#2773)
The engine mega-batch. Folds my four engine PRs and will absorb the
remaining `packages/engine` guards as commits on this branch.

**Superseded and closed:** #2722, #2741, #2766, #2770.

## Census — files converted so far

| file | before | after |
|---|---:|---:|
| `notification/notification-service.ts` | 9 | **5** |
| `runtimes/in-process-runtime.ts` | 6 | **1** |
| `eval-followups.ts` | 2 | **0** |
| `pr-comment-handler.ts` | 1 | **0** |
| `task-revert.ts` | 2 | **0** |

The last two are **census-invisible** (`Set.has(task.column)`
membership) — the class measured in #2763, which a comparison-based scan
cannot count. So the backlog number moves less than the work does,
deliberately.

## What each one actually fixed — all silent, none cosmetic

- **Notifications stopped entirely.** `handleTaskMovedAsync` compared
`data.to` to `in-review`/`done`, so on a renamed board the two
notifications operators rely on most were never sent.
- **A finished card's plan review could re-enter.** The continuation
drain's terminal test matched nothing, so a completed card's planning
continuation was handed to the executor.
- **The revert route admitted and the service refused.** The route
resolved terminal lanes; the service compared to a hardcoded pair. The
operator got a dead end from an affordance the UI and route both
offered.
- **Follow-up dedup blocked new cards forever.** A finished follow-up in
a renamed complete lane read as *open*, so the dedup matched it
permanently — defeating the intent the code documents in the line above
it.
- **The mission requeue wrote a column that may not exist**, and its
guard never matched.

## Flagged, not fixed — deliberately

- **`concurrency.ts` idle semaphore leak recovery** — the last live
caller of the running-agent predicate that does not enrich. On a renamed
board it under-counts and can reclaim a legitimately-held slot. The
enriching variant is async and this is a synchronous repair path whose
failure mode is reclaiming live work.
- **The archival `task:moved` listener** — runs on every move with no
cheap gate ahead of it; converting costs an IR resolution per move to
decide most moves are not archival.

## Notes carried from the folded PRs

Two conflicts resolved in main's favour because **main's version was
better**: `in-process-runtime`'s seam uses `terminalColumns:
ReadonlySet` (membership) where mine used `LifecycleColumns`
(first-per-role), and the test is rewritten against main's API. That
arity trap has now caught me four times, so membership is the default
shape in everything new here.

Review fixes from the folded PRs are included: the notifier's review
set, the second human-review site, the second dedup copy, the workspace
revert surface, and the file-content assertions.

## Verification

`pnpm test:gate` **GREEN** (158 + 10 + 487 + 71) · **224 passed** across
the touched engine suites · engine and dashboard `tsc` clean · `pnpm
lint` clean · census `--strict` exits 0.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 08:24:50 -07:00
gsxdsm
c53d3aec38 fix(executor): re-land the no-wip-lane fix — #2757 merged a snapshot that predated it (#2760)
## Why this exists

#2757 merged as `9a2a033b9e`, but **the third of its three fixes is not
on main**:

```
$ git show origin/main:packages/engine/src/executor.ts | grep -c wipDeclared
0
```

The merge captured my branch *before* commit `68381db72f`, so the
no-wip-lane fix was dropped while the other two landed.
`executor-execution-policy-renamed-columns` → *"a workflow with no wip
column terminalizes visibly instead of claiming the card advanced"* is
still red on main, still returning `status: null, error: null`.

This is that commit, cherry-picked cleanly onto current main. No new
work — the review discussion is in #2757.

## What it fixes (recap)

`resolveResumeLanes` defaulted `wip: lifecycle?.wip ?? "in-progress"`,
collapsing two different states:

1. **the IR failed to resolve** — defaulting is right; the `catch` arm
wants exactly this
2. **the IR resolved and declares NO wip column** — defaulting *invents
a lane the workflow does not have*

`routeGraphFailureToExecutionResume` then admitted a card resting in
that workflow's **hold** lane with incomplete steps, rehomed it,
returned `true` — and the terminalize branch never ran. Resuming into a
workflow with no implementation lane *is* "claiming the card advanced"
when nothing did.

The fix adds `wipDeclared` (declared, as opposed to defaulted) and
declines the resume when it is false — the same fail-closed rule the
sibling branch already applies with `wipColumn !== undefined` before
calling a card "already advanced". That path failed closed; this one
failed open.

IR-unavailable deliberately keeps today's behaviour (`catch` reports
`wipDeclared: true`), so an infrastructure error does not start refusing
legitimate resumes.

## Verification on current main

| check | result |
|---|---|
| the three affected files | **23 passed** |
| `pnpm test:gate` | **726** |
| engine `tsc --noEmit`, `pnpm lint` | clean |
| remove the guard | 1 failed — the fail-closed case goes red again |
| always decline | 1 failed — a legitimate resume breaks |

Full-suite blast radius was measured on #2757 before it merged: 834
files / 10,845 tests, 5 failures, both files pre-existing
(`executor-prompt`'s pause-guard 3 and `executor-abort-provenance`'s 2,
byte-identical to baseline). Zero new failures.

## Note

Worth checking whether other PRs merged in that window lost their final
commits the same way — I only noticed because I re-verified main after
the merge rather than assuming a merged PR contains what the branch
held.
2026-07-30 07:44:35 -07:00
gsxdsm
e18a6cf00c fleet: executor.ts 57 → 15 on top of #2689 — the review/wip lanes, 4 half-conversions, 8-of-19 revert proof (#2703)
**Supersedes #2691, which I am closing.** #2689 landed the terminal-pair
batch on `executor.ts` while my PR was open on the same file — we
collided, that PR won the race, and 30 of my 70 conversions are now
identical to its work. Rather than resolve 30 conflict hunks in a
20k-line lifecycle file (unreviewable, and the wrong artifact to hand
you), I rebuilt from `origin/main`.

**`executor.ts` 57 → 15.** Repo backlog 679 → **650**.

## The four that are defects, not vocabulary

**1. `isReentrantPausedAbortedInFlightNode` resolved lanes at the END,
for its return value, while its four `in-review` eligibility gates were
literals.** On a renamed board those gates all read false — so a review
card skipped the global-pause recheck, the `autoMerge === false`
refusal, the shared-branch-member arbitration **and** the
merge-confirmed refusal — and then the lane-resolved final line answered
*"re-entrant"*. FN-7214's own comment says an auto-merge-off review row
must stay terminal.

**2. The REVERSE half-conversion.**
`routeGraphFailureToExecutionResume`'s destination was already resolved
(U7's `resolveReboundColumnFor`) behind a gate that was still three
literals — so the router refused before reaching its own working move.

| direction | what happens | visible? |
|---|---|---|
| resolved gate → literal destination | card admitted, move rejected by
a board with no such column | **yes** — the move errors |
| literal gate → resolved destination | card refused; the working
recovery never runs | **no** |

Only the second is silent, which is exactly why it survived U7's own
conversion of that destination. **When you convert a destination, check
the gate in front of it in the same commit.**

**3. `routeUnusableWorktreeGraphFailureToRecovery` skipped FN-5147's
auto-merge-off gate** on a renamed board — an automatic recovery moving
a human-review-terminal card backward. #2689 converted the terminal
guard at the top of that method; this is the other half of the same
decision, which is the general risk when two people split one file.

**4. `handleGraphFailure`'s `alreadyFinalizedToReview` /
`suppressFinalizedCompletionAbort`** read `column !== "in-progress"`, so
a completed, already-finalized row looked still-in-wip: FN-6644 /
FN-6647's suppression never fired and the row was re-parked as an
operator-action pause abort — the durability gap those tickets closed.

## Two patterns worth carrying to other files

**An inert guard rarely reports "renamed board" — it reports something
that sounds like a different problem.** `finalizeAlreadyReviewedTask`
returned `"missing"` for a card sitting in review. The completion
handoff logged *"no longer active"* for a card that was actively
executing. The stuck-requeue cleanup logged *"recovered concurrently"*
about a recovery that had not happened. Three different false
explanations, one cause.

**Directions differ inside one family, so convert per method, not per
pattern.** Most wip guards read `!== "in-progress"` and REFUSE on
no-match (renamed board → silently disabled). The rerun watchdog reads
`=== "in-progress"` and SKIPS on match — there the literal never
matched, so a rerun could fire on a card **mid-execution**. A mechanical
sweep of `!== "in-progress"` fixes the refusals and leaves that
admission in place.

Also: the resolver choice inverts within a few lines. *"Is this card in
the ONE column finalize targets?"* needs the **complete** column — the
terminal union carries the legacy ids, so a card in a column merely
*named* `done` reads as already finalized and the finalize is
**skipped**. *"Is this card already finished, so do not move it?"* needs
the **union** — over-inclusion only skips a move, under-inclusion moves
a finished card out of its terminal column. Both are recorded at their
sites.

## Revert proof

19 cases in `executor-graph-failure-lanes-resolved.test.ts`, on a board
sharing **no** column id with the default lineage (on the default board
these guards are correct by coincidence — the literals *are* the board).
**8 fail on revert.** The rest are labelled **in the file** as paired
positives, default-board no-change cases, or — in one instance — a guard
that is genuinely redundant with a later lane check. I would rather
label a case as non-evidence than count it.

Two fixture corrections are recorded at their sites, both my own
assertion failing to touch the behaviour it named: asserting a router's
return value (which was already false for an unrelated reason — fixed by
spying on the recovery call), and `allowsAutoMergeProcessing` keying on
the **global** setting rather than `task.autoMerge` (fixed the fixture,
not the assertion).

## The 15 that remain, each with a reason

- **7 `to`/`from` move-effect parameters** — a move's endpoints, not a
card's resting column. Trait-hook territory.
- **2 enumeration scans** — one is a `listTasks({ column: "in-progress"
})` query whose filter cannot be converted without the query (converting
the filter alone reads as done and changes nothing); the other loops
every task, so per-task resolution is a real cost wanting a shared memo.
- **`12325`, the dependency guard** — *"is this dependency satisfied?"*
is not any single lane role. The same question exists at
`register-task-workflow-routes.ts:3995`; both should be decided once,
together.
- **`14484`** (`fromColumn === "in-review" && toColumn === "in-review"`)
— a same-lane move check that belongs with the move-effect group above.

## Verification

`pnpm test:gate` **158 / 10 / 487 / 71** · **135/135** across the 17
suites covering these paths · `tsc -p packages/engine` clean · `pnpm
lint` clean · census `--strict` exit 0, baseline re-recorded.

No changeset: `@fusion/engine` is private and the behaviour change is
confined to renamed boards.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Improved workflow execution across boards with renamed lifecycle lanes
by resolving lane targets per board instead of using fixed column names.
* Fixed review, WIP, completion, and failure-recovery behaviors to
respect the correct board snapshot (including auto-merge and terminal
work states).
* Improved artifact-recovery protection timing and tightened
execution-resume gating for failure scenarios.
* **Tests**
* Added a new lifecycle invariant test suite covering renamed-lane
recovery, resume, pause/abort, and router-gating behavior.
  * Updated lifecycle column census baseline data.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 06:48:25 -07:00
gsxdsm
78b6b5ba37 fleet: packages/engine/src/executor.ts 85 → 57 (in progress; 4 batches, plus the structural measurement this cluster needs) (#2689)
**Claiming `packages/engine/src/executor.ts`** — the largest unclaimed
cluster (self-healing.ts and scheduler.ts are taken).

## Census

| | before | after |
|---|---:|---:|
| `executor.ts` | 85 | **75** |
| repo total | 722 | **712** |
| `done` | 195 | 190 |
| `archived` | 147 | 142 |

Baseline re-recorded in the same commit; it shrinks by exactly the
converted count (10 literals across 5 sites).

## Batch 1 — terminal-lane guards

Five identical *"this card is already finished, refuse"* guards, all the
literal pair `live.column === "done" || live.column === "archived"`. On
a renamed board neither matches, so the refusal falls through — the same
inert-guard shape as #2670.

Converted to `resolveTerminalColumnsFor`, **the helper this file already
established** at line 4509 — no new abstraction. It unions the resolved
terminal columns with the legacy pair, so each converted guard is a
strict **superset** of the literal: it can refuse in more cases, never
fewer. That is what makes this batch safe without per-site behavior
review.

## The structural measurement this cluster needs

#2683 found self-healing.ts unsafe to batch because of **sync** workflow
reads — a converted guard there would resolve through a sync path that
cannot resolve a selection in production, silently falling back to
defaults. I measured whether executor.ts has the same problem, per guard
(not per line):

| context | guards |
|---|---:|
| **async** — safe, can `await resolveWorkflowIrForTask` | **71** |
| **sync** — needs threading or is not convertible in place | **14** |
| module scope | 0 |

The 14 sync-context guards are at lines 3455, 3479, 3530, 3540, 4611,
5501, 5502, 5504, 5777, 10213, 12306 (×3), 15782 — `in-progress` 4,
`in-review` 4, `archived` 3, `done` 2, `todo` 1. **I am not converting
those in place**, and I will flag rather than guess if threading
resolved data changes behavior.

So: unlike self-healing, this cluster is **83% safely convertible**,
which is why it is worth working as a batch.

## Note on #2685

Engine code converts through core's resolvers
(`resolveLifecycleColumns`, `resolveTerminalColumns`), not the dashboard
`columnRoles` helpers. So the 680-guard helper gap #2685 fixes is
**dashboard-side** — this cluster is not blocked on it.

## Verification

engine `tsc` clean · lint clean · gate green (487 + 158 + 10 + 71).

**Pre-existing failures, not caused by this change:** five tests in
`src/__tests__/reliability-interactions` fail, all in
`SelfHealingManager.recoverStarvedRefinementTriageTasks`. I confirmed by
stashing this change and re-running on a clean tree — they fail there
too. This change touches only `executor.ts` and does not go near that
path. Flagging rather than fixing: it is someone's cluster and not mine
to alter mid-flight.

## Not done

Batches 2+ (the remaining 75). I will keep working this file in this PR
with small commits, per the fleet rules.
2026-07-30 03:25:21 -07:00
gsxdsm
632d10a9b4 fix(engine): a completed-blocked guard was inert on renamed boards — plus one owner for the terminal pair (#2568)
Two commits: a behaviour-preserving extraction, then the behaviour
change.

## ⚠️ Stack note worth acting on

**#2550 and #2554 both report MERGED, but their content is not on
`main`.** They merged into their *base branches*, and the bottom of that
stack (**#2544**) is still open. Nothing in this chain has reached
`main` yet.

Nothing is lost — everything is in
`origin/feature/workflow-e2e-merge-rebound`, which is why this PR
targets it. But "merged" reads as "landed" and here it doesn't.
**Merging #2544 flows the whole chain down.**

## The bug

`parkCompletedBlockedTask` opens with *"is this card already finished?"*
and answered it with:

```ts
if (task.column === "done" || task.column === "archived") return false;
```

On a renamed board neither matches, so **the guard was inert** — and the
very next branch (`if (task.column !== "todo")`) would then have **moved
a completed card back out of its own terminal column**.

A guard that never fires does not fail a test. This one was found by
tracing the last ledger site, not by anything going red.

## Why a shared owner, not a local fix

`merger-ai`'s `isAlreadyFinalizedColumn` held the **only** copy of the
per-role terminal-pair rule — a P1 learned the hard way (PR #2471
review): a per-**set** fallback collapses to one element for a workflow
declaring `complete` but no `archived`, silently dropping the archived
half of every already-finished check.

Executor's guard was the raw literal pair, so **whoever converted it
next would have re-made exactly that mistake** — the lesson lived in a
comment in another file. Hence `resolveTerminalColumns(ir)` in core: one
owner, one place for the rule.

## Evidence, and its limits

**Commit 1 (extraction) is proven behaviour-preserving**:
`workflow-already-finalized-live-e2e` is unchanged and green through the
delegation, and the per-set mutation **still fails** through the shared
helper.

**Commit 2 (the fix) is unproven at the call site, and I'm labelling it
rather than implying otherwise.** `parkCompletedBlockedTask` is private
and reached only from inside executor dispatch — I could not drive it
end to end. So the shared helper gets its **own** tests, in both
partial-role directions, precisely because its other consumer can't
vouch for it. The call site is a one-line delegation to a tested
function.

Weaker evidence than the rest of this unit's work. Saying so, because
quietly counting it as proven is the exact failure this unit exists to
catch.

## Census

417 → 416. That ratchet (#2557) is a **ceiling**, so it stays green
without coordination; lower the pin when convenient.

## Verification

- E2E suites 10/10; helper unit tests 5/5
- core + engine `tsc --noEmit` clean
- `pnpm test:gate` green (414 + 10 + 71)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---

## Note on the conflict status (2026-07-31)

GitHub reports this PR `CONFLICTING / DIRTY`. **It is not.** Three
independent checks:

- `git rebase origin/main` on the pushed head reports *"up to date"* and
leaves the SHA unchanged — the branch is already on top of main.
- `git merge-tree` against the merge base produces **zero** conflict
markers.
- `origin/main` is unchanged at the commit this was rebased onto.

The remote SHA matches the local head, so the push landed. The
`mergeable` field is a **stale computation** — it goes stale after a
force-push and doesn't always recompute.

This branch has now been rebased and force-pushed four times against
that cached value. Worth guarding at the source: the auto-retry treats
`mergeable` as ground truth, so a stale value generates conflict notices
indefinitely. Confirming with a trial rebase or `git merge-tree` before
dispatching distinguishes "actually conflicting" from "GitHub hasn't
recomputed" — one command, and it ends the loop.

Verification on the current head: merge gate green (487 + 158 + 10),
engine + core tsc clean, lint clean, 20 tests in the affected suite,
zero unresolved threads.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 01:33:03 -07:00
gsxdsm
dca20496f4 consolidate/u7: plugins to zero + 8 executor rebound guards + resume lanes (supersedes #2607, #2635, #2640) (#2644)
Consolidation branch for U7, per the new one-branch working mode.
**Supersedes #2607, #2635, #2640** — the three of my PRs that were stuck
on review threads. My other seven (#2602, #2605, #2606, #2611, #2621,
#2628, #2633) are green with **zero unresolved threads** and are
deliberately left alone for the merge sweep.

## What is in here, file by file

| file | change | guards before → after |
|---|---|---|
| `plugins/…/glasses/src/agent-actions.ts` | gates, destinations and
degraded-resolution refusal all resolve from the task's own workflow | 2
→ 0 |
| `plugins/…/glasses/src/quick-capture.ts` | accepted capture columns
come from the board; default no longer names the deleted column | 1 → 0
|
| `plugins/…/glasses/src/settings.ts` | quick-capture default was
`triage`, the column #2515 removed | (assignment, uncounted) |
| `plugins/…/dependency-graph/src/GraphTaskNode.tsx` | redundant column
condition deleted | 1 → 0 |
| `packages/engine/src/executor.ts` | 8 rebound guards compare the
resolved column; 4 resume-eligibility literals share one resolver | 151
→ 143 (+4 off-bar) |
| `packages/engine/src/__tests__/` | 4 new suites, 26 cases | — |

`plugins/` reaches **zero** column guards with this branch.

## The three threads it closes

**#2607 — five findings, all mine, all the same rule.** I kept
*qualifying* a legacy-id fallback instead of removing it:

| attempt | rule | hole review found |
|---|---|---|
| 1 | fall back to `todo` when the role is missing | moved cards to
phantom columns |
| 2 | …only if the workflow **declares** `todo` | aliased **review**
lane named `todo` |
| 3 | …and only if no other role is assigned to it | **traitless**
parking column named `todo` |

The qualifications were the mistake. Once `resolveLanes` returns a lane
set the workflow *has* a column vocabulary, so "no column carries the
hold trait" is a complete answer — refuse. `destination()` is two lines
now, with no aliasing surface left to qualify.

Plus a sixth, which is a genuinely different state: **degraded
resolution is indistinguishable from the default board.**
`resolveWorkflowIrForTask` is total by design — a missing definition
silently returns the *default* coding IR — so a card on a custom board
whose definition could not be read resolved to `todo`/`in-progress`.
`undefined` lanes cannot express that (it means "no workflow at all",
where the legacy ids *are* the answer). The actions now refuse with 409.
#2618 would replace this check with resolver provenance; it is not
merged, so this does not depend on it.

**#2635 — "seven rebound sites remain untested."** Fair; my "same shape"
note was an assertion, not coverage. Seven of the eight need a live
graph run to reach, so the *shape* is pinned instead: a static check
that no guard in front of a rebound move compares against a column
literal, with a vacuity case (the same detection run against the
original shape) and a match-count floor (≥8), because a guard reporting
success on zero matches is worse than no guard.

**#2640 — duplicate workflow resolution.** Framed as I/O; it is also a
correctness bug. Eligibility and re-entry are two halves of one decision
and resolved the workflow separately, so a workflow edit landing between
them has the halves reading *different boards*. Now one caller-owned
memo per decision — caller-owned because a process-lifetime cache would
have to guess when a mid-flight workflow edit invalidates it.

## Behavioural findings, not tidying

- **The last-resort recovery for completed-but-stranded work did not
exist off the default lineage.** `promotedFromPlannerColumn` was false
on a renamed board, so finished work resting in planning was never
promoted; the code fell through to a review handoff that role adjacency
rejects, and the card stayed stuck with its work complete.
- **Rebound guards could not see the column their own move targeted.**
U5b converted the move target; the eight `column !== "todo"` checks in
front of it were left literal, so on a renamed board the engine moved a
card into the column it was already in — and `moveTaskInternal` runs
reset-on-entry on every real move, so at the `preserveProgress: false`
site it reset step progress a second time.
- **The FN-1404 `task:move` audit row was lying**, recording `to:
"todo"` while the move target was resolved. A run-audit trail that
disagrees with the move it describes is worse than none. Not a
comparison, so no census counts it.
- **A task interrupted by an engine pause never resumed on a renamed
board** (off-bar, `in-review`/`in-progress` literals): four comparisons
decided one question and had to agree; two of them disagreed on a
renamed board, so re-entry silently never fired.

## Revert proofs, isolated per site

| reverted | result |
|---|---|
| `destination()` back to attempt 3 | 3 of 38 fail |
| degraded-resolution refusals removed | 2 of 42 fail |
| capture set back to the legacy five | 2 of 3 fail (renamed-board
suite) |
| forward exclusions → literals | 1 of 14 fails |
| missing-wip refusal removed | 2 of 14 fail |
| `promotedFromPlannerColumn` → literals | 3 of 7 fail |
| promotion target → `"in-progress"` | 3 of 7 fail |
| one rebound guard → `!== "todo"` | 1 of 3 fails (static shape) |
| resume lanes → legacy trio | 1 of 5 fails |

Every conversion is paired with a negative — a forward move, a
not-a-planner-lane card, a default-lineage card, an unresolvable
workflow — so neither "always fire" nor "never fire" can pass for
"resolve the role".

## Commit discipline

Twelve commits, each one thing: the code move (`resolvePlannerLanes` out
of `triage.ts`) is separate from every behavior change, and each review
fix is its own commit with its own revert proof.

## Verification

- `pnpm test:gate` **71/71**
- 162/162 across the glasses plugin's 19 files; 26/26 across the four
new engine suites
- engine + glasses typecheck clean; `pnpm lint` clean

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Engine recovery and retries now work correctly with renamed or
customized workflow columns.
  * Tasks in manual-intake columns are no longer automatically planned.
* Agent actions and quick capture now respect each board’s declared
columns and lifecycle stages.
* Awaiting-approval tasks are recognized regardless of their current
column.
* Command Center SDLC funnel stages now accurately reflect customized
workflows.

* **Documentation**
* Added guidance for safely changing workflow-column logic and
interpreting lifecycle-column checks.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 00:52:55 -07:00
gsxdsm
6ed284f36a drop the dead semaphore parameter from dropPreHeldExecutorSlot (#2574)
Small follow-on to the cross-project cap removal.

`dropPreHeldExecutorSlot(taskId, semaphore?)` released a cross-project
semaphore slot. That semaphore is deleted, and **all 16 production call
sites passed `this.options.semaphore`**, which nothing wires any more —
so the release was a no-op on an always-undefined value: an optional
parameter that reads as if it does something.

## What is *not* deleted

Pre-held slots are **dual-purpose**: a cross-project semaphore slot
**and** the FN-8453 per-project coordinator reservation. Only the first
is gone. The reservation is the half that matters — every rejection path
funnels through this helper so an early scheduler/triage return cannot
permanently consume a project slot — and it stays. That is why this is a
parameter change, not a helper deletion.

Sites that still hold a semaphore reference release it **explicitly**
next to their drop, so behaviour is unchanged for any caller that
supplies one. Nothing wires one in production today, but silently
leaking a slot for a caller that does is not a trade a cleanup is
allowed to make.

## One real leak fixed — found by a failing test, not by reading

`ProjectAdmissionCoordinator.admitOldest`’s release lambda took the
pre-held branch and **returned**, relying on the deleted parameter to
hand the host slot back. With the parameter gone, that branch unwound
the registration and the reservation while **leaking the host slot** the
attempt had acquired. The release is now unconditional across both
branches.

Worth noting how it surfaced: the test that caught it (`drops a declined
candidate’s pre-held executor slot`) asserted `semaphore.activeCount`,
which I had initially assumed was just coupling to the deleted half. It
was not — it was pinning a real invariant.

## Tests

Five cases in `concurrency.test.ts` pinned `sem.activeCount` through a
drop. Each is re-pointed at the surviving contract — registration and
reservation unwound, nothing left for a later pass to “take” — with the
semaphore assertions moved to the sites that now own the release.

## Verification

`pnpm lint` clean · engine `tsc` clean · `pnpm test:gate` green ·
`concurrency.test.ts` **56/56**. The 8 `triage.test.ts` failures are
**pre-existing** — reproduced identically with this branch’s `triage.ts`
replaced by main’s.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 23:34:52 -07:00
gsxdsm
3aa942ee5f capacity: spawned agents count against the project agent count (#2579)
Two configurable numbers per project. `maxSpawnedAgentsPerParent` (5)
and `maxSpawnedAgentsGlobal` (20) were a **third and fourth** limiter
with private budgets invisible to both.

## This closes a hole, not just knobs

A spawned child **is** an agent and gets **its own git worktree**
(branched from the parent’s — the tool’s own description says so), but
children were counted by **neither** capacity gate. A fan-out could put
up to 20 extra worktrees on disk while the scheduler believed the
project was at its configured limit. The operator’s two numbers were
simply wrong about what was running.

## The old caps also measured the wrong thing

`totalSpawnedCount` decrements on child cleanup, but the per-parent
**set** is cleared only when the **parent task** ends. So
`maxSpawnedAgentsPerParent` throttled *cumulative* spawns across a
task’s life rather than *concurrent* ones — a long-running task could
exhaust its budget with five children that had all long since finished,
and the operator had no way to see why.

## Fix

`fn_spawn_agent` gates on the same project agent count every other lane
uses (`computeTopLevelConcurrencyClaimedFromStore`) plus live children.
One number, one answer, no private budget that can disagree with the
board.

The refusal names **Max Concurrent Tasks** — a control the operator
actually has. The old messages pointed at settings that no longer exist,
which is worse than no message: it sends someone hunting for a knob that
is not there.

## Verification

**Revert-proof, measured:** restoring the private budgets turns **3 of
the 4** new cases red — a project at 1/1 could still spawn, which is
precisely the hole. `executor.ts` restored byte-identical.

`pnpm lint` clean · core + engine `tsc` clean · `pnpm test:gate` green
(414 + 10 + 71) · new suite 4/4 · `settings-default-descriptions` 4/4.

There was no spawn-capacity test before this; the file is new.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Spawned agents now count toward the project’s **Max Concurrent Tasks**
capacity.
* Agent spawning is blocked when capacity is reached, including
concurrent spawn attempts.
* **Bug Fixes**
  * Prevented over-allocation during simultaneous agent spawns.
  * Restored available capacity when agent creation fails.
* **Changes**
  * Removed separate per-parent and global spawned-agent limits.
  * Updated settings to reflect the revised capacity controls.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 23:34:41 -07:00
gsxdsm
31e49b684a TAKING default-workflow-hooks.ts + executor.ts + live-agent-count.ts + 6 dashboard files: reopen semantics by role, and the census's blind spot in both directions (13 sites) (#2628)
Batched conversion of every lifecycle-column guard I hold, plus the
three the census could not see. **Six files to zero, repo-wide 60 → 49
by a comment-stripped unanchored sweep.** Each conversion has an
isolated revert proof and a paired negative case, and the one code move
is a separate commit from the behavior changes.

## Per-file before → after

Counts from a comment-stripped, unanchored `(===|!==) ["']triage["']`
sweep over `packages/*/src` + `plugins/*/src`, excluding tests.

| file | before | after | note |
|---|---:|---:|---|
| `core/default-workflow-hooks.ts` | 4 | **0** | |
| `core/task-store/moves.ts` | 5 | **4** | only the flag-ON mirror
converted; the flag-OFF inline block is the parity reference and stays |
| `engine/executor.ts` | 3 | **0** | **absent from the 45-guard list** —
see below |
| `core/live-agent-count.ts` | 2 | **0** | duplication removed; answer
deliberately unchanged |
| `engine/replan-target.ts` | 2 | **0** | both were comment prose, not
guards |
| `core/agent-prompts.ts` | 3 | **0** | ROLE comparisons, never column
guards |
| `engine/usage-limit-detector.ts` | 2 | **0** | ROLE comparisons |
| `dashboard/app/components/DocumentsView.tsx` | 1 | **0** | real column
guard |
| `dashboard/app/components/TaskChatTab.tsx` | 2 | **0** | ROLE |
| `dashboard/app/components/AgentLogViewer.tsx` | 1 | **0** | ROLE |
| `dashboard/app/components/effective-model-resolution.ts` | 1 | **0** |
ROLE |
| `dashboard/app/hooks/useTasks.ts` | 1 | **0** | ROLE |
| `dashboard/…/command-center/MissionControlPanel.tsx` | 1 | 1 | alias
table, marked `DELIBERATE-LITERAL` with its reason |

## The census errs in BOTH directions

This is the finding I would most like carried into the remaining work.

- It **flagged 10 sites that were never column guards.** `role ===
"triage"` / `agentType === "triage"` compare an **AGENT ROLE**. The
planner *lane* is named `triage` and keeps that name — U11 removed the
*column*. Worse than noise: the obvious "finish the migration" edit is
to rename the role, and that silently empties the planner's prompt
template and mis-binds its model markers. `PLANNER_AGENT_ROLE` now names
it, so the two vocabularies are distinguishable by grep and a rename
fails loudly (revert proof: 4 tests, two of them pre-existing).
- It **missed 3 real guards in `executor.ts`**, because the pattern
matches `column`/`toColumn`/`fromColumn` and those locals are named
`from` and `originColumn`. A census keyed on variable names will keep
missing guards wherever a local was named for its role in the function.

## Two real defects, not tidying

**1. A renamed board could merge with its re-review never run.**
`default-workflow-hooks.ts` is named for the default workflow, but the
store runs it on the flag-ON path for *every* workflow — the trait
registry resolves hooks by trait id, not by workflow. Its reopen
predicates listed the default lineage's column names, so on a renamed
board **no reopen effect fired at all**. One of them clears
`workflowStepResults`, which `getTaskMergeBlocker` reads: a card bounced
out of review carried its old `passed` result back in, and that
satisfies the merge gate. Same regression the graph-owned-crossing
carve-out exists to prevent, arriving through the other door. (Two
smaller ones rode along: failure state never cleared on a renamed
reopen, and an operator dragging a card back to the queue never parked
it, so the scheduler re-dispatched what they had just pulled back.)

**I forgot the carve-out on my first pass, and that was worse than not
converting.** A role-resolved clear plus a *name*-matched exemption
means a renamed board takes the clear and never the exemption,
destroying the remediation input the graph had just written. My own
paired negative test caught it.

**2. The last-resort recovery for completed-but-stranded work did not
exist off the default lineage.** In `recoverCompletedTask`,
`promotedFromPlannerColumn` was false on a renamed board, so finished
work resting in the planning lane was never promoted — the code fell
through to `handoffTaskToReview` straight from the planning column, and
role adjacency has no planning → review edge, so the handoff was
rejected and the card stayed stuck with its work complete. I converted
the promotion **target** too: resolving the lane and then moving to a
literal `in-progress` is the half-conversion I have already been burned
by twice this program, where the guard starts admitting cards and the
move then sends them to a column the board does not declare.

## E2E evidence

`renamed-board-reopen.pg.test.ts` drives a **real PostgreSQL store** and
a real `moveTask` on a workflow whose columns carry the standard traits
under non-default names. The unit tests cannot show this: if `moves.ts`
passed `undefined`, every unit case still passes via the no-basis
fallback while the real board keeps the old behavior. **Proof it is
load-bearing: forcing `moveLifecycleColumns` to `undefined` fails 2 of
3.** The executor suite covers both the split-role and the MERGED
post-U11 shape.

## Revert proofs, isolated per site

| change reverted | result |
|---|---|
| reopen predicate → literal names | 4 of 10 fail |
| reopen field clears → literal names | 2 of 10 fail |
| `userPaused` hold lane → literal `todo` | 1 of 10 fail |
| graph carve-out → literal names | 1 of 10 fail |
| store passes `undefined` lifecycle columns | 2 of 3 fail (real PG) |
| `promotedFromPlannerColumn` → literals | 3 of 7 fail |
| two-hop condition → `=== "triage"` | 1 of 7 fails |
| promotion target → `"in-progress"` | 3 of 7 fail |
| `isPlannerColumnFor` → literals | 1 of 7 fails |
| live-agent-count: one arm dropped | 2 of 11 fail |
| DocumentsView: trait branch removed | 3 of 7 fail |
| planner role renamed to `"planner"` | 4 fail (2 pre-existing) |

Every conversion is paired with a negative case (a forward move, a
not-a-planner-lane card, a default-lineage card, a renamed column with
no traits), so neither "always fire" nor "never fire" can pass for
"resolve the role".

## Deliberately NOT converted, with reasons

- **`moves.ts` flag-OFF inline block (4).** That branch *is* the legacy
path, kept verbatim so the two can be parity-checked. Converting it
erases the reference implementation.
- **`live-agent-count.ts`'s no-flags fallback.** Reachable, and there is
nothing to resolve from — `enrich…FromFlags` exists for callers with
board flags rather than an IR, so a column missing from that map is the
renamed case. "Not intake" is as much a guess as "todo is intake", and
Running/Waiting are complements, so a card matching neither arm is
reported as neither and the footer's queued total under-reports it. The
real fix is at the caller; four new cases pin that flags override the
legacy answer **in both directions**. What did change is the
duplication: two hand-written copies of one rule now call one named
function.
- **`MissionControlPanel`'s `FUNNEL_STAGES`.** An alias table of column
*names* where `triage` sits beside `signal` and `backlog`. Command
Center aggregates across projects, so there is no single workflow to
resolve traits from — the honest conversion is a data change, not a
predicate change.
- **`DocumentsView` with no traits.** Same no-basis rule; the documents
list is full of historical columns absent from the current board. A case
asserts a renamed column with no traits still reads as "working",
documenting the gap rather than hiding it.

## Fixture findings

Each cost a red run that looked like the code under test:

- a `merge-blocker` column needs a reachable merge-class node, or
`parseWorkflowIr` rejects the workflow;
- a back-edge must be `kind: "rework"`, and a rework edge is legal only
**into** a node with `config.reworkRegion: true`;
- a workflow gets role-level transitions only when it declares wip +
review + complete + **archived** plus a planning lane — without the
archived column, adjacency falls back to order-derived neighbours and
`checking -> queued` is not a legal move at all;
- `recoverCompletedTask` only *reaches* the promotion seam when nothing
is left to gate; without passed `plan-review`/`code-review` rows it
re-enters the workflow graph and returns first, so a naive fixture
silently tests the wrong branch and every assertion reads "no moves
happened" for an unrelated reason.

## Verification

- `pnpm test:gate` **71/71**
- new suites: 10/10 reopen-semantics, 3/3 renamed-board-reopen (real
PG), 7/7 executor-planner-lanes, 7/7 documents-status-dot, 4/4
planner-role-is-not-a-column
- neighbours: 132 + 10 + 482 (gate shards), 350/351 engine
planning/replan suites, 64/64 agent-prompts, 51/51 usage-limit-detector,
11/11 live-agent-count, 11/11 dashboard hook/log suites
- the single engine failure (`executor-fast-mode-workflows.test.ts` ›
"raw fast mode still invokes non-executable review seam nodes")
**reproduces with my changes stashed** — pre-existing on `origin/main`
- typechecks clean for core, engine, and dashboard-app
(`tsconfig.app.json`; `tsconfig.json` checks nothing under `app/`);
`pnpm lint` clean

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 22:39:14 -07:00