Commit Graph

2243 Commits

Author SHA1 Message Date
gsxdsm
529ab26ba4 FN-8695: fix dashboard assertion CRUD validation
Restore dashboard assertion status edits and deletes for store-generated IDs.

- Validate assertion route IDs against the MissionStore-generated format while preserving legacy IDs.
- Dispatch assertion deletes through the shared confirmation handler and surface request failures.
- Cover current ID CRUD routes and desktop/mobile deletion behavior.

Files changed:
 .changeset/fn-8695-assertion-id-validation.md      |   7 ++
 docs/missions.md                                   |   2 +-
 .../dashboard/app/components/MissionManager.tsx    |  51 ++++++++---
 .../MissionManager.delete-confirm.test.tsx         |  99 ++++++++++++++++++++
 .../mission-assertion-id-validation.test.ts        | 100 +++++++++++++++++++++
 packages/dashboard/src/mission-routes.ts           |   8 +-
 6 files changed, 249 insertions(+), 18 deletions(-)

Fusion-Task-Id: FN-8695

Fusion-Task-Lineage: b7e2d70f-d3c7-4fdd-92f6-989cf414570b

Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
2026-08-01 13:10:26 -07:00
gsxdsm
5d06c92b77 FN-8713: handle missing auth-status query data
Keep auth-status available when direct callers omit credential-instance query data.

- Default absent request query data to an empty selection.
- Preserve validation and prevent dangling instances from falling back to other credentials.
- Add coverage for omitted, valid, and dangling credential-instance queries.
- Add a patch changeset for the auth-status fix.

Files changed:
 .changeset/fn-8713-auth-status-query.md            |  7 +++
 .../register-auth-routes-hermes-additive.test.ts   | 67 ++++++++++++++++++++--
 .../dashboard/src/routes/register-auth-routes.ts   |  9 ++-
 3 files changed, 77 insertions(+), 6 deletions(-)

Fusion-Task-Id: FN-8713

Fusion-Task-Lineage: 0bcb965e-0606-4c7b-aefc-e7e3461ae75d

Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
2026-08-01 12:53:07 -07:00
gsxdsm
2ab81d954b FN-8712: fix GitLab delete listener logging
Route GitLab delete-close listener failures through the structured dashboard logger.

- Replace console warnings with a named structured logger
- Record normalized error details alongside task and lifecycle metadata
- Cover listener-boundary failures without duplicate closure or unhandled rejections

Files changed:
 .../src/__tests__/gitlab-delete-close.test.ts      | 46 ++++++++++++++++++----
 packages/dashboard/src/gitlab-delete-close.ts      | 18 ++++++---
 2 files changed, 51 insertions(+), 13 deletions(-)

Fusion-Task-Id: FN-8712

Fusion-Task-Lineage: 466a5b80-efa7-4961-bdb6-79277f6deb41

Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
2026-08-01 10:45:13 -07:00
gsxdsm
79d2a73a10 FN-8690: document Grok CLI provenance blocker
Document the unavailable Grok CLI source provenance and preserve the existing usage gate.

- Record provenance investigation results and the BLOCKED-NO-SOURCE hand-off
- Validate FN-8690 evidence sections and canonical verdicts
- Explain why the API-supplied percentage gate remains unchanged

Files changed:
 docs/solutions/integration-issues/grok-cli-usage-data-source.md | 64 ++++++++++++++++++++++
 packages/dashboard/src/__tests__/grok-usage-finding-doc.test.ts | 21 +++++++
 packages/dashboard/src/usage.ts                                  |  3 +
 3 files changed, 88 insertions(+)

Fusion-Task-Id: FN-8690

Fusion-Task-Lineage: 4521063d-2a9a-4710-9609-35b7582a2f2d

Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
2026-08-01 08:52:08 -07:00
gsxdsm
006cc40454 FN-8685: add durable cross-process task deletion consumers
Deliver durable, replay-safe cross-process task deletion observation.

- Add PostgreSQL lifecycle consumer cursors, leases, acknowledgements, retention, and recovery.
- Start named consumers in dashboard, serve, and engine runtime paths.
- Preserve delete integration metadata while suppressing replayed GitHub and GitLab side effects.
- Cover outbox identity, observed delivery, fencing, and reconciliation behavior.

Files changed:
 ...fn-8685-cross-process-task-deleted-observers.md |   7 +
 .../fn-8685-task-deleted-outbox-consumers.md       |   7 +
 docs/architecture.md                               |   8 +-
 ...tgres-cross-process-task-deleted-observation.md |   8 +-
 docs/storage.md                                    |  10 +-
 packages/cli/src/commands/dashboard.ts             |   9 +-
 packages/cli/src/commands/serve.ts                 |   9 +-
 packages/cli/src/project-context.ts                |   9 +-
 .../task-deleted-outbox-consumer.pg.test.ts        | 157 ++++++++
 ...-deleted-observed-dispatch-side-effects.test.ts |  36 ++
 .../task-lifecycle-consumer-identity.test.ts       |  22 ++
 packages/core/src/index.ts                         |  11 +
 .../0041_fn_8685_task_lifecycle_consumers.sql      |  88 +++++
 packages/core/src/postgres/schema-applier.ts       |  16 +-
 packages/core/src/postgres/schema/project.ts       |  45 +++
 packages/core/src/postgres/startup-factory.ts      |   4 +
 packages/core/src/store.ts                         |  54 ++-
 .../__tests__/lifecycle-outbox-writer.test.ts      |   4 +-
 .../core/src/task-store/archive-lifecycle-2.ts     |   1 +
 packages/core/src/task-store/lifecycle-ops.ts      |  13 +-
 packages/core/src/task-store/lifecycle-outbox.ts   |   2 +
 packages/core/src/task-store/project-store-ops.ts  |   4 +-
 .../src/task-store/task-deleted-outbox-consumer.ts | 333 +++++++++++++++++
 .../task-store/task-lifecycle-consumer-identity.ts |  32 ++
 .../task-store/task-lifecycle-consumer-registry.ts | 396 +++++++++++++++++++++
 .../task-store/task-lifecycle-event-retention.ts   | 104 ++++++
 packages/core/src/task-store/task-mutation-ops.ts  |   1 +
 packages/dashboard/src/github-tracking-state.ts    |  12 +-
 packages/dashboard/src/gitlab-delete-close.ts      |   3 +
 packages/dashboard/src/gitlab-split-close.ts       |   7 +-
 packages/dashboard/src/project-store-resolver.ts   |   9 +-
 packages/engine/src/project-manager.ts             |   4 +-
 packages/engine/src/project-runtime.ts             |   2 +-
 packages/engine/src/runtimes/in-process-runtime.ts |  17 +-
 packages/engine/src/self-healing.ts                |  27 ++
 35 files changed, 1439 insertions(+), 32 deletions(-)

Fusion-Task-Id: FN-8685
Fusion-Task-Lineage: 63eca9ac-d2af-44b0-ba79-388a950148d3
Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
2026-08-01 06:09:04 -07:00
gsxdsm
4b306f10bd FN-8689: document Grok CLI source provenance gap
Record the unrecoverable Grok CLI provenance chain and preserve the unmeterable usage state.

- Document installed asset identity, attempted provenance retrievals, and the static-blocked verdict.
- Clarify that the legacy billing request is not verified CLI /usage behavior.
- Add a regression test for the provenance finding and credential-safe documentation.

Files changed:
 docs/solutions/integration-issues/grok-cli-usage-data-source.md | 172 +++++++++++----------
 packages/dashboard/src/__tests__/grok-usage-finding-doc.test.ts |  45 ++++++
 packages/dashboard/src/usage.ts                                 |   8 +-
 3 files changed, 143 insertions(+), 82 deletions(-)

Fusion-Task-Id: FN-8689

Fusion-Task-Lineage: 473fc008-2e08-47f6-8f53-152da5b2c31a

Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
2026-08-01 04:41:50 -07:00
gsxdsm
a7591853eb FN-8688: document unverified Grok CLI usage source
Document the provenance gap that prevents deriving Grok CLI usage data safely.

- Record version-skewed source and sanitized billing-replay evidence
- Keep absent Grok usage fields authenticated but unmeterable pending source-backed confirmation
- Link the usage-provider rationale to the investigation record

Files changed:
 .../grok-cli-usage-data-source.md                  | 93 ++++++++++++++++++++++
 packages/dashboard/src/usage.ts                    |  4 +-
 2 files changed, 95 insertions(+), 2 deletions(-)

Fusion-Task-Id: FN-8688

Fusion-Task-Lineage: 5aafa505-968f-4f63-83db-35e623f84052

Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
2026-08-01 04:19:46 -07:00
gsxdsm
c978cdbacf FN-8687: close imported GitLab issues on task deletion
Add GitLab close-on-delete behavior for imported source issues.

- Close linked GitLab issues by default while honoring leave and delete action semantics.
- Preserve split-close ownership, skip merge requests, and limit malformed tracking fallback to deletion.
- Document the lifecycle contract and add dashboard coverage.

Files changed:
 .changeset/fn-8687-gitlab-close-on-delete.md       |   7 ++
 docs/gitlab-parity-inventory.md                    |  20 +++-
 docs/settings-reference.md                         |   1 +
 docs/task-management.md                            |   5 +-
 .../src/__tests__/gitlab-delete-close.test.ts      | 101 +++++++++++++++++++++
 .../src/__tests__/gitlab-lifecycle.test.ts         |  35 +++++++
 .../gitlab-parity-inventory-documentation.test.ts  |   7 +-
 .../__tests__/gitlab-source-issue-close.test.ts    |   8 ++
 .../src/__tests__/gitlab-split-close.test.ts       |  12 +++
 packages/dashboard/src/gitlab-delete-close.ts      |  90 ++++++++++++++++++
 packages/dashboard/src/gitlab-lifecycle.ts         |  33 +++++--
 packages/dashboard/src/gitlab-split-close.ts       |   2 +-
 packages/dashboard/src/index.ts                    |   1 +
 .../dashboard/src/routes/register-git-github.ts    |   7 ++
 14 files changed, 313 insertions(+), 16 deletions(-)

Fusion-Task-Id: FN-8687

Fusion-Task-Lineage: 48bc9159-fc35-43cb-b72e-1b6260d9e3cb

Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
2026-08-01 04:08:57 -07:00
gsxdsm
0e8f6769ba FN-8666: add credential instance pickers
Add credential-instance selection across dashboard model configuration surfaces.

- Expose provider credential instances in custom model dropdowns and task, workflow, and settings forms.
- Persist per-role credential-instance overrides through dashboard APIs and effective model resolution.
- Document credential-instance precedence and cover dropdown/settings behavior, including task state synchronization.

Files changed:
 .changeset/fn-8666-credential-instance-picker.md   |   7 ++
 docs/settings-reference.md                         |   2 +
 docs/workflow-steps.md                             |   2 +-
 packages/dashboard/app/api.ts                      |   1 +
 packages/dashboard/app/api/models-usage.ts         |  27 ++++-
 packages/dashboard/app/api/tasks.ts                |   8 ++
 .../app/components/CustomModelDropdown.css         |  12 ++-
 .../app/components/CustomModelDropdown.tsx         |  71 ++++++++++++-
 .../dashboard/app/components/InlineCreateCard.tsx  |  23 +++-
 packages/dashboard/app/components/ListView.tsx     |  29 +++++-
 .../app/components/ModelSelectionModal.tsx         |  25 ++++-
 .../dashboard/app/components/ModelSelectorTab.tsx  |  29 +++++-
 packages/dashboard/app/components/NewTaskModal.tsx |  23 +++-
 .../dashboard/app/components/QuickEntryBox.tsx     |  28 +++++
 .../dashboard/app/components/SettingsModal.tsx     |   2 +
 .../dashboard/app/components/TaskDetailModal.tsx   |  32 +++++-
 packages/dashboard/app/components/TaskForm.tsx     |  18 ++++
 .../app/components/WorkflowNodeEditor.tsx          |  13 ++-
 .../app/components/WorkflowSettingsPanel.tsx       |  38 ++++++-
 ...ustomModelDropdown.credential-instance.test.tsx | 111 ++++++++++++++++++++
 .../__tests__/WorkflowSettingsPanel.test.tsx       |   9 +-
 .../app/components/effective-model-resolution.ts   |   8 +-
 .../settings/sections/GlobalModelsSection.tsx      |  31 +++++-
 .../settings/sections/ProjectModelsSection.tsx     | 116 ++++++++++++++++-----
 .../ProjectModelsSection.chatDefault.test.tsx      |  33 +++++-
 packages/dashboard/app/hooks/useFavorites.ts       |   6 +-
 packages/dashboard/app/hooks/useModelsCache.ts     |   5 +-
 .../src/routes/register-task-workflow-routes.ts    |  17 ++-
 28 files changed, 655 insertions(+), 71 deletions(-)

Fusion-Task-Id: FN-8666

Fusion-Task-Lineage: 6c3711b3-68b8-47a0-ac36-4b64846adf39

Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
2026-08-01 04:03:26 -07:00
gsxdsm
23c9992e7b FN-8682: add GitLab split-close issue comments
GitLab source issues now receive an explanatory split handoff before closure.

- Add a GitLab split-close service that posts normalized child-task notes before closing source issues.
- Wire split-close handling across project stores while preserving merge-request and failure safeguards.
- Document GitLab lifecycle parity and cover split-close behavior with tests.

Files changed:
 .changeset/fn-8682-gitlab-split-close-comment.md   |   7 ++
 docs/gitlab-parity-inventory.md                    |   2 +-
 docs/task-management.md                            |   1 +
 .../gitlab-parity-inventory-documentation.test.ts  |   2 +
 .../src/__tests__/gitlab-split-close.test.ts       | 113 +++++++++++++++++++++
 packages/dashboard/src/gitlab-lifecycle.ts         |   6 +-
 packages/dashboard/src/gitlab-split-close.ts       |  98 ++++++++++++++++++
 packages/dashboard/src/gitlab-tracking-state.ts    |   6 +-
 packages/dashboard/src/index.ts                    |   1 +
 .../dashboard/src/routes/register-git-github.ts    |   6 ++
 10 files changed, 237 insertions(+), 5 deletions(-)

Fusion-Task-Id: FN-8682

Fusion-Task-Lineage: 8eca9888-b10a-4ea9-b9d3-fc09a39022ab

Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
2026-08-01 03:19:00 -07:00
gsxdsm
3793b576d8 FN-8673: comment on split source issue closures
Explain GitHub source-issue closures when imported work is split into subtasks.

- carry split closure context through task deletion and triage
- post one explanatory comment before closing source and tracking GitHub issues
- preserve exactly-one-comment behavior when close retries after transient failures
- document split closure behavior and cover source/tracking scenarios

Files changed:
 .changeset/fn-8673-split-close-issue-comment.md    |  7 ++
 docs/settings-reference.md                         |  2 +-
 docs/task-management.md                            |  1 +
 .../task-delete-caller-attribution.test.ts         | 32 +++++++
 packages/core/src/index.ts                         |  2 +-
 packages/core/src/store.ts                         | 10 +--
 packages/core/src/task-delete-attribution.ts       | 16 ++++
 .../core/src/task-store/archive-lifecycle-2.ts     | 14 ++--
 packages/core/src/task-store/archive-lifecycle.ts  |  6 +-
 packages/core/src/types.ts                         | 13 +++
 .../src/__tests__/github-tracking-state.test.ts    | 98 +++++++++++++++++++++-
 packages/dashboard/src/github-tracking-state.ts    | 82 +++++++++++++++---
 packages/engine/src/__tests__/triage.test.ts       |  4 +
 packages/engine/src/triage.ts                      | 10 +++
 14 files changed, 268 insertions(+), 29 deletions(-)

Fusion-Task-Id: FN-8673

Fusion-Task-Lineage: 904c2445-64d9-47b9-b706-f64b23c4e3a6

Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
2026-08-01 02:47:25 -07:00
gsxdsm
8a6949dd24 FN-8661: resolve selected credential instances for sessions
Resolve requested provider credential instances before creating agent sessions.

- Thread lane credential instance selections through planning, validation, execution, review, and merge sessions.
- Resolve selected instances into runtime credential stores while retaining provider-default fallback behavior.
- Preserve selected credentials for mission validation, executor retries, and spawned child agents.

Files changed:
 .../fn-8661-credential-instance-resolution.md      |  7 ++
 AGENTS.md                                          |  1 +
 docs/architecture.md                               |  2 +-
 docs/secrets.md                                    |  2 +
 docs/settings-reference.md                         |  1 +
 .../dashboard/src/__tests__/routes-auth.test.ts    | 80 +++++++++++++++++++
 .../dashboard/src/routes/register-model-routes.ts  | 78 +++++++++++++++++++
 .../src/__tests__/agent-session-helpers.test.ts    | 16 ++++
 .../credential-instance-resolution.test.ts         | 49 ++++++++++++
 packages/engine/src/agent-heartbeat.ts             |  1 +
 packages/engine/src/agent-runtime.ts               |  9 ++-
 packages/engine/src/agent-session-helpers.ts       | 67 +++++++++++-----
 packages/engine/src/auth-storage.ts                | 90 ++++++++++++++++++----
 packages/engine/src/executor.ts                    | 29 ++++++-
 packages/engine/src/merger-ai.ts                   |  2 +
 packages/engine/src/merger.ts                      |  5 ++
 packages/engine/src/mission-execution-loop.ts      |  4 +-
 packages/engine/src/pi.ts                          |  7 +-
 packages/engine/src/pr-response-run-ops.ts         |  1 +
 packages/engine/src/reviewer.ts                    |  7 ++
 packages/engine/src/triage.ts                      |  2 +
 21 files changed, 420 insertions(+), 40 deletions(-)

Fusion-Task-Id: FN-8661

Fusion-Task-Lineage: 1e34a3ce-0857-4619-9746-ce0dc12dc2ba

Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
2026-08-01 02:05:03 -07:00
gsxdsm
5f12044168 FN-8652: add multiple provider credential instances
Enable operators to create, select, and remove named credentials for each supported provider.

- Add provider-auth instance discovery and credential mutation endpoints.
- Add dashboard authentication controls, status handling, and instance coverage.
- Document instance behavior and add a release changeset.

Files changed:
 .changeset/fn-8652-provider-credential-instances.md       |   7 +
 docs/secrets.md                                    |   2 +
 docs/settings-reference.md                         |   4 +
 packages/dashboard/app/api.ts                      |   1 +
 packages/dashboard/app/api/provider-status.ts      |  94 +++++-
 .../dashboard/app/components/SettingsModal.tsx     | 138 ++++-----
 .../AuthenticationSection.instances.test.tsx       |  75 +++++
 .../settings/sections/AuthenticationSection.css    |  13 +
 .../settings/sections/AuthenticationSection.tsx    | 199 +++++++-----
 .../dashboard/src/__tests__/routes-auth.test.ts    |  27 +-
 packages/dashboard/src/routes.ts                   |  12 +-
 .../dashboard/src/routes/register-auth-routes.ts   | 334 ++++++++++++++++++---
 .../src/__tests__/provider-auth-instances.test.ts  |  65 ++++
 packages/engine/src/provider-auth.ts               |  89 ++++++
 14 files changed, 845 insertions(+), 215 deletions(-)

Fusion-Task-Id: FN-8652

Fusion-Task-Lineage: 39568d58-7f57-4d17-97cb-1837323b3a94

Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
2026-08-01 01:07:59 -07:00
gsxdsm
5596d915ab FN-8647: quarantine flaky Kimi K3 catalog test
Quarantine the timing-sensitive Kimi K3 SDK catalog test without changing timeout budgets.

- Reuse the native model registry once per test file.
- Add the observed CI timeout to the dashboard quarantine ledger and config.
- Document validation and timeout-budget preservation requirements.

Files changed:
 docs/testing.md                                    |  8 ++++++++
 ...ister-model-routes-kimi-k3-supplemental.test.ts | 23 ++++++++++++++++++++--
 packages/dashboard/vitest.config.ts                |  8 ++++++++
 scripts/lib/test-quarantine.json                   |  5 +++++
 4 files changed, 42 insertions(+), 2 deletions(-)

Fusion-Task-Id: FN-8647

Fusion-Task-Lineage: 31e79677-d923-4003-a8e8-082159334e65

Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
2026-07-31 22:32:18 -07:00
gsxdsm
b65bc0bbb4 FN-8650: preserve Grok CLI authentication status
Keep unavailable Grok CLI billing data distinct from authentication expiry.

- Preserve billing outcomes for usage windows, missing data, HTTP responses, and transport failures.
- Mark CLI authentication expired only after observed 401 or 403 responses.
- Cover unmeterable authenticated accounts in usage and indicator tests.
- Add a patch changeset for the corrected indicator behavior.

Files changed:
 .changeset/fn-8650-grok-cli-auth-expired.md        |  7 ++
 .../components/__tests__/UsageIndicator.test.tsx   | 16 +++++
 packages/dashboard/src/__tests__/usage.test.ts     | 79 ++++++++++++++++++++--
 packages/dashboard/src/usage.ts                    | 57 ++++++++++++----
 4 files changed, 139 insertions(+), 20 deletions(-)

Fusion-Task-Id: FN-8650

Fusion-Task-Lineage: f18cceb6-1092-4073-b540-9c7a5434fbcd

Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
2026-07-31 19:20:16 -07:00
gsxdsm
f31a716a2a FN-8629: prevent false Grok usage percentages
Prevent omitted Grok billing percentages from being displayed as fully consumed credits.

- Require a finite CLI-supplied credit usage percentage before creating a billing window.
- Cover omitted, zero, invalid, and non-weekly Grok billing responses.
- Add a patch changeset for the corrected usage display.

Files changed:
 .changeset/fn-8629-grok-usage-percent.md       |  7 +++
 packages/dashboard/src/__tests__/usage.test.ts | 71 ++++++++++++++++++++++++--
 packages/dashboard/src/usage.ts                | 13 ++---
 3 files changed, 77 insertions(+), 14 deletions(-)

Fusion-Task-Id: FN-8629

Fusion-Task-Lineage: b5b7c83b-e34f-43d9-a31a-d1fd769c4eb8

Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
2026-07-31 13:43:28 -07:00
gsxdsm
476c5c360c test(dashboard): pin the Reliability endpoint's three lane reads (extract seam + pin) (#3237)
## What

Pins the **Reliability endpoint's three lane reads** — the last
uncovered resolver cluster the repo-wide audit found.

Two commits, deliberately separate:
1. **refactor** — extract the three resolves behind
`resolveReliabilityLanes(store)`. Behaviour-preserving, no test changes.
2. **test** — pin all three through that seam.

## Why a seam was needed

The three resolves lived inline in the `/api/health/reliability` route
closure. Blinding any of them left the **entire dashboard suite green —
21,582 tests** — and the only way to reach them was booting
`createServer` behind a mock-the-world shell the slow-test rule forbids.

**And the obvious test would not have helped.**
`reliability-metrics.test.ts` exercises `countEntriesInto`,
`countBouncesOut` and `inReviewDurationMetrics` with lane sets **passed
in by hand**. That proves the collaborators honour a resolved set; it
says nothing about whether the caller passes one. *A unit test of the
collaborator can never fail when the caller's resolve is blinded* — the
same trap the audit note records for `reads.ts`, where a suite written
for the exact conversion still could not see it.

The seam is the caller. It resolves, so blinding a resolve fails a test
of it.

## Measured — each blind fails exactly its own case

| blinded | fails |
|---|---|
| `REVIEW_ROLES` | "resolves the board's OWN review lane" |
| `["countsTowardWip"]` | "resolves the board's OWN wip lane" |
| `["complete"]` | "resolves the board's OWN complete lane" |

```
converted: Tests 6 passed (6)
each blind: Tests 1 failed  (its own case only)
reliability-metrics.test.ts + this file: 28 passed
typecheck clean; lint clean; fnxc-future-dates: none added
```

That isolation is the point: **three resolves in one function invite a
copy-paste that hands the same set to all three**, and every positive
assertion would still pass. There is a paired negative asserting each
renamed lane appears in *its* bucket and nowhere else — without it the
duration metric could silently measure review → review.

Also pinned: the degrade path. An unreadable workflow list must not fail
the endpoint, so the legacy ids still answer.

## What breaks without the conversion

On a board that renames either lane, every underlying query returns `{}`
— so `tasksEnteredInReview` and `tasksBouncedToInProgress` are zero for
every day, and `inReviewFailureRate7d` divides one zero by another and
reports a **healthy** rate. It produces a NUMBER, not an error, and the
number says everything is fine. An operator reading 0% review failures
beside a populated audit list has no reason to suspect the metric is
blind.

## The one observable difference in the refactor, stated not buried

The complete-lane read moves from *after* the counting `Promise.all`
into the same phase as the review/wip pair. These are pure reads of
workflow definitions — no writes, no ordering dependency — so the
resolved values are identical; only the concurrency shape changes (three
parallel reads instead of two-then-one). Flagging it because
"behaviour-preserving" should be a claim someone can check, not an
assertion.

## Audit status

With this, **3 of the 4 flagged sites are closed**. Remaining:
`cli/commands/task.ts:660`, where the glyph decision is inline in
`runTaskList` and the same seam argument applies — but its sibling test
file already documents that driving that function needs the forbidden
shell, and extracting a helper there would produce a test that *looks*
like coverage while leaving the resolve unpinned. Left flagged rather
than faked.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

- **New Features**
- Reliability health metrics now recognize configured review,
work-in-progress, and completion lanes, including renamed workflow
lanes.

- **Bug Fixes**
- Improved fallback behavior when workflow definitions are unavailable,
preserving compatibility with legacy lane configurations.
- Ensured lane resolution remains isolated by role for more accurate
reliability metrics.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-31 13:23:20 -07:00
gsxdsm
623581837a fix(engine): mock provider sends 0-based steps — test mode full-task runs complete again (#3231)
Found by a live browser E2E of the coding workflow in test mode: every
scripted full-task run failed at `steps#0:step-execute` with `Step 4 out
of range (task has 4 steps)`, rebounding through recovery forever.

**Root cause:** `fn_task_update.step` has been **0-based since FN-6607**
(executor.ts FNXC:StepNumbering — the old `step - 1` conversion made
Step 0 impossible to mark). `mock-provider.ts` still sent `index + 1`,
so test mode marked steps 1..N instead of 0..N-1: Step 0 (Preflight)
never completed and step N threw out-of-range. Test mode's full-task
path has been broken since June.

**Also fixes the test that pinned the bug:** `mock-provider.test.ts`
expected `{ step: 1 }` for a fixture whose first unfinished step is
index 0 — the expectation encoded the 1-based off-by-one.

Verified: 12/12 mock-provider tests; the live E2E instance completes the
task after this patch (see follow-up screenshot in the session).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-31 12:54:58 -07:00
gsxdsm
d6079970e8 fix(self-healing): 18 recovery rebounds hardcoded todo and THREW on a renamed board (#3150, first slice) (#3152)
First slice of #3150. `self-healing.ts` held **26** `moveTask` calls
with a legacy literal target; this converts the **18 `todo` rebounds**.

## Why this is worse than a guard, and documented already

`task-store/moves.ts` records it from a previous incident:

> `moveTaskInternal` **REJECTS** a target the workflow does not declare
(`TransitionRejectionError: unknown-column`) … completion handoff did
not silently no-op — it **THREW**.

Every one of these 18 is a **recovery**. On a renamed board they threw
instead of rebounding, so the strand each sweep exists to clear survived
*and* the sweep reported failure. The reliability layer meant to be the
backstop was the layer that broke.

## Why the census never saw it

It counts **comparisons** against legacy ids. A move target is an
**argument**. That is the third blind spot of the same instrument, and
all three have now produced real defects found by hand:

| blind spot | found this session |
|---|---|
| definitions | `GITHUB_TRACKING_EDITABLE_COLUMNS` — tracking
unreachable on renamed boards (#3149) |
| collections | swept: 30 sites, 29 already correct, 1 defect (the
above) |
| **targets** | **this** — 26 in one file, 31 tree-wide |

## Why 18 sites at once is safe

`resolveReboundTargetForTask` **degrades to `"todo"`** when no workflow
resolves, and `self-healing.ts` already used it at line 745. On every
board we ship, the resolved answer *is* `todo` — so default behaviour is
unchanged **by construction**, not by inspection. The control case pins
exactly that, and it is the reason this can land as one change rather
than eighteen.

## Scope, and what I deliberately did not touch

Converted: the 18 `todo` rebounds.

**Not** converted: the `done`, `archived` and `in-review` targets. They
need different helpers and genuine reasoning about which lane a
completion or an archive belongs in — converting them by analogy is
exactly the half-conversion this program keeps paying for. Sites with no
resolver in scope are unchanged.

The audit behind the split is in the commit: of 26 sites, 5 had resolved
lanes in scope, 4 had an IR, 17 had nothing — and `lanesOfReclaim`
returns **Sets**, which is the wrong arity for a target (a move takes
exactly one column, per the `moves.ts` note).

## Verification

| | result |
|---|---|
| engine `tsc` | **0 errors** |
| **all 43 self-healing suites** | **843 passed** |
| census `--strict` | exit 0, **unchanged** — invisible to it |
| `check-inert-sync-lanes` | exit 0 |
| differential | restoring the literal → **1 failed \| 1 passed**,
renamed case only |

The new test drives a **public entry point**
(`reconcileInReviewUnmetDependencies`, the FN-6793 contract) rather than
calling the helper directly, so it covers the producer path too.

One harness note worth keeping: the first version of the test failed
**upstream** of the target, because the sweep selects rows via
`resolveProjectColumnsForRoles` — a *project-level* resolver reading
`listWorkflowDefinitions`, not the task's own selection. Without that
mocked, the renamed card was never considered and the failure looked
like the fix not working. That distinction (project-level vocabulary vs
per-task IR) will bite the next slices too.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Tasks now move to workflow-specific rebound, completion, and archive
columns instead of fixed default destinations.
* Retrying and recovering tasks works correctly on boards with renamed
lifecycle columns.
* Added safe fallback behavior for workflows without custom lifecycle
settings.
* **Tests**
* Added coverage to prevent legacy hardcoded task destinations and
verify renamed-column recovery scenarios.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 10:24:22 -07:00
gsxdsm
6bc90ccbe2 fix(core): allow-list the legacy workflow IR — it found a fourth bug my grep missed (#3185)
## The name is the defect

`BUILTIN_CODING_WORKFLOW_IR` reads like the default and **is** the
legacy workflow (`builtin:legacy-coding`). Post-U11 they differ by
exactly one column — `triage` — the one a caller most often wants
absent.

**Four bugs have come from reaching for it by name:**

1. two move-path resolvers disagreed on the no-selection default →
*"workflow move policy preflight is stale"* on every flag-on move
(recorded in `resolveDefaultWorkflowIr`'s own header)
2. the TUI board rendered a `triage` lane the default board lacks —
#3178
3. `deleteWorkflow` re-homed occupants into `triage` — #3183
4. **`board-workflows.ts`** described a *custom* workflow whose
definition failed to load using legacy columns — the #3178 symptom
through the dashboard route. **Fixed here.**

It type-checks, it is the obvious identifier, and on the five shared
columns it behaves correctly. The mistake only shows on the column that
differs.

## I said the sweep was complete last round. It wasn't.

My grep excluded paths and truncated at `head -10`; it missed two sites.
**The allow-list found both on its first run.**

That is the lesson the sibling sync-resolver ratchet already records —
*"FOUND BY THIS RATCHET, not by the grep that seeded the list"* — and I
had just quoted that file while repeating the mistake.

## One site is allow-listed rather than fixed, and I tried the fix first

`workflow-graph-executor.run()`'s default `ir` is unreachable in
production (both callers pass it explicitly). But
`workflow-graph-executor-parity.test.ts`, in the **engine-core gate
suite**, drives the method *without* the argument to assert the
historical seam sequence.

Switching it to the catalog default rewrites what "parity" means:
**measured, 6 gate tests fail** with `expected 'failure' to be
'success'`. Reverted, and recorded at the call site *and* in the
allow-list entry so nobody repeats the experiment.

That is what an allow-list is for: a legitimate narrow use next to a
plausible-looking wrong one.

## Guard construction

Follows the repo's existing call-site allow-lists (sync resolver, engine
blocking-shellout, detached-spawn script guard).

- **Comments stripped before scanning** — `activity-analytics.ts` and
`TaskContextMenu.tsx` name this constant in notes *about past bugs*
while correctly avoiding it. Counting prose would train readers to
allow-list mentions.
- **Anti-vacuity**: the scan still sees the catalog's own uses, so a
renamed constant or broken walker cannot make the guard pass by finding
nothing.
- **Stale-entry**: the list cannot rot into files that no longer touch
it — the decay every ledger in this repo has hit.

## Measured

- Guard **3/3**; `tsc --noEmit` clean in core, engine, dashboard.
- census `--strict`, `check-fnxc-future-dates` clean.

## Census

**No movement — that is the point.** This class has no column literal to
count, which is why the census never saw any of the four bugs.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 08:45:31 -07:00
gsxdsm
c220455e3a fleet: 10 inline fallback arms become named sets (census 102 → 92) (#3061)
## Census

| | column guards |
|---|---|
| before | **102** |
| after | **92** |

Five files drop to **0** guards each. Baseline re-recorded in the same
commit.

## A cluster the census could not distinguish from real debt

**Every site here is already converted.** Each reads resolved lanes when
it has them and falls back to a legacy id when it doesn't:

```ts
reviewColumns ? reviewColumns.has(task.column) : task.column === "in-review"
```

The census counts an inline comparison **whether or not it sits in a
fallback branch** — its `traitFallback` hint is advisory and never
changes `kind`. So ten correctly-converted guards sat on the backlog
permanently, and the number stopped distinguishing *work still to do*
from *documented degraded answers*.

Naming the fallback set fixes the bookkeeping without touching
behaviour: `new Set(["in-review"]).has(x)` answers exactly what `x ===
"in-review"` answered.

## Files

| file | sites | what they gate |
|---|---|---|
| `restart-recovery-coordinator.ts` | 4 | three shared review gates +
one `??` default |
| `github-tracking-state.ts` | 2 | complete / archived lane predicates |
| `planner-overseer.ts` | 2 | wip / review classification |
| `async-mission-store-queries.ts` | 2 | terminal complete / archived |
| `register-task-workflow-routes.ts` | 2 | wip promotion target,
archived respecify guard |

**No behaviour change is claimed and none is intended** — that's the
point. These were already right; only the accounting was wrong.

## Worth the fleet's attention

Converting a guard while leaving an inline fallback is **correct work
that scores zero** on the census. My own first pass at `reads.ts` did
exactly that — behaviourally correct, census unmoved. Anyone converting
this way is doing real work the number won't credit, and the backlog
will look stuck.

## Measured

| check | result |
|---|---|
| engine suites | **173 tests green** |
| core mission suites | **70 tests green** |
| dashboard route suites | **211 tests green** |
| five gates + strict census | green |
| `tsc` (core, engine, dashboard) | clean |

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Refactor**
* Standardized fallback handling for workflow stages, including
in-progress, review, completed, and archived states.
* Preserved existing behavior when explicit workflow column settings are
available or unavailable.
* Improved consistency across task tracking, planning, and recovery
workflows.

* **Chores**
* Updated lifecycle tracking baselines to reflect current source-file
coverage.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 03:34:27 -07:00
Drew Donaldson
c1c1b964af fix(dashboard): expose full permission-mapped task toolset in chat sessions (#2376)
## Bug
Chat-session tool surface missing task-mutation tools that exist outside
of chat, even when the agent's permission record grants them.\n\nRepro:
agent-09dcf8b2 (role: custom, CEO) in NextGenEHS has tasks:archive /
tasks:delete / tasks:merge / tasks:retry / tasks:update true in its
permission record with permissionPolicy.presetId = unrestricted and
task_agent_mutation = allow. Calling fn_task_archive / fn_task_delete /
fn_task_merge in chat returns: Tool fn_task_* not found.\n\nRoot cause:
packages/dashboard/src/chat.ts createChatFusionToolset() built a
hardcoded narrow chat-only allowlist while heartbeat registered the
complete lifecycle surface unconditionally.\n\nFix:\n- Add exported
factories in packages/engine/src/agent-tools.ts for missing lifecycle
tools: fn_task_archive, fn_task_unarchive, fn_task_delete,
fn_task_retry, fn_task_pause, fn_task_unpause, fn_task_duplicate,
fn_task_merge, fn_task_update, fn_task_add_dep, fn_task_promote,
fn_trait_list, fn_ask_question, fn_reflect_on_performance,
fn_read_evaluations, fn_update_identity, fn_send_message,
fn_read_messages.\n- Wire those factories into
createChatFusionToolset(). Mission/ideation mutations stay behind
missionMutationGated. Agent-scoped tools still require agentId.\n-
Re-export from packages/engine/src/index.ts.\n- Regression test:
packages/dashboard/src/__tests__/chat-toolset-permissions.test.ts (3/3
passing). Existing chat.test.ts (14/14 passing).

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Chat now exposes task lifecycle actions—including archive, retry,
pause, duplicate, merge, and dependency updates—when permitted by the
agent’s action controls.
- Added support for identity updates and evaluation viewing in
agent-linked chats.
- Existing read-only tools remain available, while restricted actions
stay hidden when authorization is unavailable.
- **Tests**
- Added regression coverage for authorized and unauthorized chat tool
surfaces, including preservation of read-only capabilities.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: gsxdsm <gsxdsm@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 21:51:09 -07:00
gsxdsm
189f237a07 fix(security): make the approval and permission gates actually enforce (#2457)
## Why

An AI agent autonomously deleted a live task (FN-8609) via
`fn_task_delete`. Investigating that led to two audits, which found the
gate machinery is well-built and its *enforcement* well-tested — but the
*decision boundary* was unguarded and the tools that matter never
reached the gate at all.

The headline defect: all ~76 tools in `packages/cli/src/extension.ts`
are delivered through `hostExtensionPaths` → `DefaultResourceLoader`,
which never passes through `wrapToolsWithActionGate`. `fn_task_delete`
*was* classified as gated in `gating-classifications.ts` — that
classification was dead code. Roughly 15 destructive tools were ungated
even for an operator who had locked their policy all the way down.

Second: `POST /api/approvals/:id/decision` took the approver identity
from the request body, validated for shape only, with no check that the
decider was human or differed from the requester. Combined with a bash
tool whose command string was never inspected (only `params.cwd`), an
agent could read the daemon token from `~/.fusion/settings.json`, find
its own pending request, and approve it.

## Scope

This does **not** tighten defaults. The `unrestricted` preset is
byte-for-byte unchanged — `git diff` on `agent-permission-policy.ts` is
empty — and regression tests assert that an out-of-the-box install
behaves exactly as before. The bug was never "the default is too
permissive"; it was "strict policy doesn't enforce." This makes turning
security up actually work.

The one deliberate exception: the containment that stops an agent
escalating its *own* privileges (reading the daemon token / credentials,
calling the approvals API to self-approve) applies at every preset
including `unrestricted`. That is a privilege-escalation boundary rather
than a permission preference — if it only engaged under strict policy it
would not have prevented the incident that prompted this.

## What changed

8 bisectable commits:

- **Approval lifecycle** — self-approval blocked via server-derived
deciders; same-verdict replay 409s; decide re-reads and re-validates
inside the transaction; expiry TTLs; `markCompleted` ownership check;
session identity registry in core.
- **Engine gates enforce for real** — unclassified tools resolve to a
policy-governed category instead of hardcoded `allow`; missing-policy
fail-open closed; bash containment floor + exact-command approval
binding.
- **Dashboard decision routes** — stop trusting client-supplied actors
(decision, bypass-review, worktrunk → 403 on forged actors).
- **`fn serve` authenticated by default** — auto-mints a token following
the existing `fn dashboard` precedent; `--no-auth` opts out.
- **Sibling entry points closed** — user-sourced hard-cancel moves, ACP
execute-once approvals, plugin task-store gating.
- **pi-extension principal resolution** — the extension resolves the
acting principal and can withhold or policy-gate the previously ungated
destructive tools.
- **Root-cause bonus fix** — `findLatestByDedupeKey` was broken in
PostgreSQL backend mode (already-parsed jsonb fed through a string-only
parser), so approved-grant redemption **never matched in production**,
minting duplicate requests. This explains the live DB state of 17
approved / 0 completed. *(Also cherry-picked to `main` as `a9b30013bb`,
since it is an active production defect on its own.)*
- **Review follow-ups** (`627f1b1fa8`) — operator-configured
provisioning privilege and a configurable grant TTL; see below.

## Review follow-ups

**Provisioning privilege is operator-configured, not role-derived.**
`isCallerPrivileged` had gone from `caller.reportsTo == null` (every
top-level agent privileged — permanent escalation by creating a
manager-less agent) to `caller.role === "ceo"`, which swapped an
implicit rule for a magic string: any agent config can claim that role,
while an operator who genuinely wants a privileged agent had no
supported way to say so. Privilege now derives solely from
`agentProvisioning.trustedAgentIds` / `trustedRoles` and fails closed
when settings are unresolvable.

It is also no longer forwarded to `resolveAgentProvisioningPolicy` as
`isPrivileged`, because that flag short-circuits ahead of
`alwaysApproveDelete` — a trusted caller was bypassing delete approval
entirely. The policy applies the same trusted rules itself, in the right
order. The function now governs only the org-chart escape hatch (acting
outside your own direct reports).

**Grant TTL defaults to 1 hour and is configurable.** Approval →
redemption is not instantaneous: an operator approving from their phone,
an engine restart, a queued lane, or a task waiting on a worktree all
routinely exceeded 15 minutes, after which the grant expired and the
agent silently re-requested. One hour remains far short of the
"redeemable forever" hazard the TTL exists to bound. Override via
`FUSION_APPROVAL_GRANT_TTL_MS` or `configureApprovalRequestTtls()`;
invalid overrides are ignored rather than widening the window to
infinity or collapsing it to zero.

## Behavior changes requiring operator review before rollout

1. `fn serve` requires a bearer token by default (`--no-auth` opts out);
unauthenticated clients get 401.
2. Agents can no longer run withheld destructive tools
(`fn_task_delete`, `fn_task_bypass_review`,
mission/milestone/slice/feature/workflow deletes, `experiment_finalize`,
`skills_install`). Operators keep them via CLI/dashboard. **This is the
incident fix.**
3. Agents get provisioning privilege only when the operator lists them
in `agentProvisioning.trustedAgentIds` / `trustedRoles`; the
provisioning gate is now live in production. Previously-implicit
privilege (top-level position, or a `ceo` role) no longer grants
anything on its own.
4. Decision replay 409s (was 200); pending approvals expire after 24h,
approved grants after 1h (configurable); bash approvals bind per exact
command.
5. Forged/body actors on decision, bypass-review, worktrunk routes →
403; `archive-all-done` requires `{confirm:true}` (external scripts
affected).
6. `fn_secret_get` approvals grant exactly one reveal (previously
granted nothing and looped forever); ACP approvals are execute-once
(previously infinite reuse).
7. Bash containment denies token/credential/approvals-API commands in
all agent sessions at every preset.

## Verification

Independently re-run against the branch, not just self-reported:

- 5 typechecks (core, engine, cli, dashboard `tsconfig.json` +
`tsconfig.app.json`) — clean
- `pnpm lint` — clean
- `pnpm test:gate` — 379 passed
- `pnpm build --force` — green (a plain `pnpm build` skips packages as
unchanged and does **not** compile the branch)
- `pnpm check:changesets` — clean
- ~650 file-scoped tests including new negative-path suites for the
decision boundary, which previously had **zero** test coverage

`packages/engine/src/__tests__/plugin-runner.test.ts` fails 56/80 —
**verified pre-existing**, reproducing identically at base commit
`93a403af67` on `main`. Not in the merge gate.

### A mutation check that failed to fail

Worth recording, because it nearly shipped an untested security fix. The
first mutation check on the provisioning change reintroduced the `ceo`
hardcode and **all 17 tests still passed** — the tests asserted through
the policy path, which can no longer observe `isCallerPrivileged` at
all, precisely because `isPrivileged` is no longer forwarded there.
Org-chart cases that do exercise the function were added; the hardcode
now fails exactly 1 of 19, and restoring is green. A green mutation run
is only meaningful if the test can actually see the code under test.

## Known limitations (stated, not papered over)

- The bash containment floor is string-matching: a cost-raiser, not a
sandbox. Quoting, encoding, `$HOME`, symlinks, or an interpreter
one-liner can evade it. The durable protection is the decision route
refusing agent-originated deciders — the filter is the belt, not the
braces.
- Approval expiry is lazy (evaluated at decide/complete/redeem), not
swept, so an expired pending row stays visible in lists until touched.
- The extension's require-approval path returns a pending message but
cannot suspend a pi session mid-turn; engine-side pause hooks cover
engine lanes only.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Security**
* Hardened approval and permission gating with server-side decider
attribution, self-approval blocking, ownership checks, replay/race
protection, and status/TTL enforcement.
* Added fail-closed behavior for sensitive/unclassified tools and
sandbox provisioning approvals.
* Blocked credential/approval access via bash containment; plugin
destructive task operations now require explicit permission.
* **New Features**
* `fn serve` now defaults to bearer-token auth, with `--no-auth` as the
explicit opt-out.
* **Bug Fixes**
* Improved task move-source attribution (`moveSource: "user"`) and
tightened dashboard archive/bypass confirmation and operator attribution
behavior.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-30 21:50:37 -07:00
gsxdsm
fd795883c5 feat(missions): per-mission taskPrefix override for triaged task ids (#2347)
## Summary
Maintainer re-land of
[#2334](https://github.com/Runfusion/Fusion/pull/2334) (fork
`flexi767:feat/per-mission-task-prefix`) after resolving merge conflicts
with current `main`.

Fork push was unavailable despite `maintainerCanModify`, so this branch
carries the conflict resolution.

### Feature
- Optional per-mission `taskPrefix` for triaged task ids (inherits
project prefix when unset)
- Dashboard MissionManager + routes + store/triage plumbing
- Postgres migration for `project.missions.task_prefix`

### Conflict resolution
- Main claimed migration **0026** (bigint counters) and **0027**
(workflow IR pin)
- Mission task-prefix migration renumbered **0026 → 0028**
- Baseline `0000_initial.sql` includes `task_prefix` on missions
- `legacy.ts` keeps code-org re-exports; `missions.ts` carries
`taskPrefix` on create/update types

## Test plan
- [ ] CI green (lint/typecheck/build/gate)
- [ ] Create mission with custom prefix; triage feature → task ids use
that prefix
- [ ] Clear mission prefix via PATCH null; new tasks inherit project
prefix

Closes / supersedes #2334 once this lands (or re-point the fork PR).

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Missions can now set an optional per-mission task ID prefix
(overriding the project default).
* Added task prefix support to mission create/edit UI and dashboard
APIs, including normalized uppercase values and validation.
* **Bug Fixes**
* Improved commit hook generation for custom prefixes and special
characters, with safer shell handling to prevent unsafe interpretation.
* **Chores**
* Added PostgreSQL migration and schema-applier support to persist and
propagate mission task prefixes, including upgrade/backfill coverage.
* **Tests**
* Added backend and UI/API test coverage for task-prefix creation,
clearing, and ID minting behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-30 21:35:23 -07:00
gsxdsm
60bfebdc98 fix(reliability): the duration query hid its lane ids inside a SQL template (#2875)
The Reliability panel's **third and last** blind input — and my own
loose end. #2861 fixed the two counts beside it, so the panel went from
uniformly wrong to **partially** wrong: entries and bounces populated,
duration reporting `no-in-review-entries` forever. Partial blindness is
harder to notice than total, which is why finishing it matters more than
one site suggests.

```sql
metadata->>'to' = 'in-review'
  OR (metadata->>'from' = 'in-review' AND metadata->>'to' = 'done')
```

## The class, not just the site

**This shape is invisible to every check we have.** The lifecycle census
scans `===`/`!==` comparisons; the unwired-lane-parameter guard scans
declarations. Neither sees a lane id inside a `sql` template, so this
class is **not in the backlog total at all** — the number is a floor for
this reason as well as the usual one.

`scripts/check-sql-column-literals.mjs` (#2841, in flight) is the
detector for exactly this: it freezes the surface at 30 sites rather
than converting any, so this one was unowned. That PR and this one are
complementary — it stops the surface growing, this shrinks it by one.

## The fix

Lanes resolve **once per call** via `resolveProjectColumnsForRoles` and
arrive as parameterised equality fragments, one branch per id — no
interpolated list, no string building.

Resolution lives in `getInReviewDurationEventsImpl` because that is
where the store is; `async-audit.ts` takes a bare `db` handle and cannot
resolve anything. Best-effort, defaulting to the legacy pair, so a
caller that cannot resolve keeps exactly today's query.

**The union is correct rather than a widening hack**, for the same
reason as #2861: these are *move records*, and a past move recorded the
column name as it was at the time. A board renamed last month has rows
under both ids, so the honest query covers both — which is precisely
what `resolveProjectColumnsForRoles` returns.

## Tested against real PostgreSQL, deliberately

This is a **SQL predicate** change. A mocked store would assert the
arguments and prove nothing about the query that actually runs — which
is the entire risk when the literal lives inside `sql`. The new case
inserts real `activity_log` rows on a renamed board and reads them back
through the real store method.

The legacy-lane case in the same file stays green, which is the
compatibility half.

**Revert proof, measured:** restore the hardcoded fragments and the new
case fails with

```
expected [] to deeply equal [ 'renamed-entered', 'renamed-done' ]
```

## Verification

- `pnpm test:gate` — 161 / 487 / 13 / 71 passed
- `pnpm lint` — clean
- `tsc --noEmit` (`@fusion/core`) — clean
- `activity-log-parity.pg.test.ts` — 5 passed against real PostgreSQL

With this, all three Reliability inputs read the board's own lanes.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Reliability duration metrics now work correctly with renamed workflow
lanes.
* Completion tracking recognizes configured completion lanes instead of
relying on fixed defaults.
* Improved handling of transitions between multiple review lanes and
review-to-work-in-progress movements.
* Legacy lane behavior remains supported when configured lane
information is unavailable.

* **Tests**
* Added coverage for renamed lanes, historical lane IDs, and transition
edge cases.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 17:59:34 -07:00
gsxdsm
defe48d30f fix(core): per-workflow metrics read zero on a renamed board (#2866)
Second of the 14 lane-bound SQL sites from #2839, after #2864.
Independent of it — different file, different caller argument.

## The defect

`aggregateWorkflowAnalytics` filtered in SQL on `t."column" = 'done'`
and `IN ('in-progress','in-review')`. On a renamed board those match
nothing, so `tasksCompleted`, `tasksInProgress` and `tasksInReview` come
back **zero for every workflow** while the board is busy. Nothing
errors.

Same shape and same fix as #2864: resolve per **project** via
`resolveProjectColumnsForRoles`, bind an `IN` list, and thread the store
from the single Command Center caller so the parameter has a supplier
immediately rather than becoming an inert seam.

## What the test caught that I had not

**The renamed case still failed with the query fixed.**

The bucketing at lines 296–297 already uses `isWipColumnRole` /
`isReviewColumnRole` — correctly converted — but those read
`query.columnFlagsByName`, which production supplies and my fixture did
not. So:

- the **SQL** decides *which rows come back*;
- the **trait map** decides *which bucket each row lands in*.

Both halves have to be right. Fixing only the query would have shipped a
"conversion" that still reported zero on a renamed board, and the file
would have scored as converted twice over. That is exactly the
partial-conversion shape this program keeps re-finding — caught here
only because the test asserts `tasksInReview` alongside
`tasksCompleted`, since those two paths take **different** resolved sets
(complete vs wip+human-review). Asserting the completed count alone
would have left the second conversion unproven.

## Measured

Reverted, only the renamed case flips:

```
✓ default vocabulary: completed and in-review work are counted
× renamed vocabulary: completed and in-review work are counted
✓ renamed vocabulary: a card in the HOLD lane counts as neither
✓ without a lane store, the legacy ids still answer
  Tests  1 failed | 3 passed (4)
```

The hold-lane negative is there so resolving real lanes cannot degrade
into "every column counts" — trading an undercount for an overcount is
harder to notice than the original bug.

## Scope

The sync SQLite arm in the same file keeps its literals: it throws in
backend mode and has no production caller, the same dead-arm conclusion
reached for `cleanupStaleMergeQueueRowsImpl` on #2839.

## Verification

`pnpm test:gate` green · both Command Center analytics suites 8/8 ·
`tsc` core 0, dashboard 0 · lint 0 · changeset included.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 17:40:48 -07:00
gsxdsm
890e1f87e7 fix(core): issue panels reported nothing fixed on a renamed board (#2871)
Fourth and last of the lane-bound analytics sites from #2839, after
#2864, #2866 and #2870.

## The defect

`aggregateGithubIssueAnalytics` and its GitLab twin filtered their
resolved-issue query on `"column" = 'done'`. On a renamed board that
matches nothing, so `fixed` is **zero**, the resolved-issue list is
empty, and `net` reports every filed issue as still outstanding — while
the team closes issues all week. Nothing errors.

Same fix as the previous three: resolve per **project** via
`resolveProjectColumnsForRoles`, bind an `IN` list, thread the store
from each Command Center caller so the parameters have suppliers
immediately.

## Both providers in one change, deliberately

These two files are **copies** — same query, only the provider literal
differs — and a copy is exactly what gets half-fixed. Converting one and
not the other type-checks, passes that provider's test, and leaves the
second silently broken with no signal anywhere. The suite runs every
case against both, so the pair cannot drift.

## Measured

Reverted, exactly the two renamed cases fail — **one per provider** —
while both default-vocabulary controls, both WIP-lane negatives, and
both omitted-store legacy cases stay green:

```
✓ github: default vocabulary counts a resolved issue
× github: renamed vocabulary counts a resolved issue
✓ github: renamed vocabulary does NOT count an issue still in the WIP lane
✓ github: without a lane store, the legacy id still answers
✓ gitlab: default vocabulary counts a resolved issue
× gitlab: renamed vocabulary counts a resolved issue
✓ gitlab: renamed vocabulary does NOT count an issue still in the WIP lane
✓ gitlab: without a lane store, the legacy id still answers
  Tests  2 failed | 6 passed (8)
```

That the failures are symmetric is itself the check on the copy-paste
risk.

## Scope

The sync SQLite arms keep their literals: they throw in backend mode and
have no production caller, the same dead-arm conclusion as
`cleanupStaleMergeQueueRowsImpl` on #2839.

## Verification

`pnpm test:gate` green · Command Center + GitLab issue analytics suites
10/10 · `tsc` core 0, dashboard 0 · lint 0 · changeset included.

---

**This closes the lane-bound half of #2839.** All 14 sites the
hand-review identified as genuinely vocabulary-bound are now converted
across four PRs. What remains there is the 11 `!= 'archived'`
exclusions, which are probably correct as literals — archiving writes
`task.column = 'archived'` unconditionally as a state rather than a lane
— plus one dead SQLite arm. Those need per-site judgment, not
conversion.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 16:53:12 -07:00
gsxdsm
216632bd3a fix(core): task-duration stats were computed from an empty set on a renamed board (#2870)
Third of the 14 lane-bound SQL sites from #2839, after #2864 and #2866.
Independent of both.

## The defect

`aggregateProductivityAnalytics` filtered its duration query on
`"column" = 'done'`. On a renamed board that matches nothing, so the
entire task-duration distribution — median, p90, average, total — is
computed from an **empty row set** and reports zeros while the project
ships work. Nothing errors.

Same shape and fix as the previous two: resolve per **project** via
`resolveProjectColumnsForRoles`, bind an `IN` list, thread the store
from the single Command Center caller so the parameter has a supplier
immediately rather than becoming an inert seam.

## Measured

Reverted, only the renamed case flips:

```
✓ default vocabulary: a finished task contributes to the duration stats
× renamed vocabulary: a task in the RENAMED complete lane contributes
✓ renamed vocabulary: a task still in the WIP lane does NOT contribute
✓ without a lane store, the legacy id still answers
  Tests  1 failed | 3 passed (4)
```

## The negative asserts the median, not just the count

This fix's failure mode is **worse than the bug it fixes**. Resolving
too many lanes would pull unfinished work into the distribution and
produce a plausible-but-wrong median — a number nobody questions — where
the bug produces an obvious zero. So the WIP-lane case asserts
`medianMs` is null as well as `completedTasks` being 0.

## A fixture error worth naming

My first version asserted `taskDuration.count`. `TaskDurationSummary`
exposes `completedTasks`. Every case failed with `expected undefined to
be 1` — **including the controls** — which reads exactly like a broken
product until you notice the control is failing too. A control that
fails is a fixture bug, not a finding; that asymmetry is the fastest way
to tell them apart.

## Scope

The sync SQLite arm keeps its literal: it throws in backend mode and has
no production caller, the same dead-arm conclusion as
`cleanupStaleMergeQueueRowsImpl` on #2839.

## Verification

`pnpm test:gate` green · both Command Center analytics suites 8/8 ·
`tsc` core 0, dashboard 0 · lint 0 · changeset included.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 16:50:00 -07:00
gsxdsm
b85e5f90e1 fix(create): two task-CREATE destinations named a lane the board does not have (#2843)
Both files sat at **census-zero** and both wrote real cards into columns
no workflow declares. The census scores `===` comparisons, so a lane
literal passed as a **call argument** is invisible to it — one of the
four census-blind classes. These are the only two explicit-`column`
creates in production:

```
packages/dashboard/src/routes/register-gitlab.ts:108   column: "triage"
packages/cli/src/extension.ts:5243                     column: "todo"
```

## The two defects

**`register-gitlab.ts` — `column: "triage"`, a column U11 DELETED.**
This one is broken on *every* board, not only renamed ones: the default
lineage is now `todo | in-progress | in-review | done | archived`.
`createTask` already resolves the intake column of the workflow it
selects (`resolvedEntryColumn`), and an explicit `column` **overrides**
that resolution — which is why the stale literal survived U11. Nothing
rejects the write and nothing logs it: the route answers `201` with a
task id and the imported card is simply not on the board. Same shape as
the `task-update.ts` triage defect fixed earlier in this program.

Fix: omit `column` and let `createTask` resolve intake.

**`extension.ts` `fn_delegate_task` — `column: "todo"`.** On a workflow
whose ready lane is named anything else, the delegated card goes to an
undeclared column: written, reported to the caller as delegated, never
visible to the agent it was delegated to.

Fix: resolve the selected workflow's `hold` lane. **Deliberately not**
"omit the column like the GitLab route" — the tool's own contract is
*"the task goes to the ready-to-work lane and the target agent picks it
up on its next heartbeat"*, so inheriting intake resolution would park a
delegated card in a manual-intake lane waiting for a human. That would
be a behaviour change; `hold` is the role that names the lane the
literal meant.

## New helper: `resolveWorkflowColumnForRole(store, role, workflowId?)`

The **write**-shaped counterpart to `resolveProjectColumnsForRoles`. The
read helper unions in the legacy ids because an extra id in a query set
is inert; here the same trick is a silent wrong write (post-U12 an
undeclared column is a `TransitionRejectionError` on move, a phantom
lane on create), so it returns one column from one workflow, or
`undefined`.

**A contract I got wrong twice, now pinned by a test.** `undefined`
means *"this workflow declares no such column"* and nothing else.
`resolveWorkflowIrById` never throws and never returns nothing — an
unregistered builtin id, a missing definition row and a failing read all
resolve to the default coding IR (branded via `markFellBack`). So an
unreadable workflow yields the **built-in** hold lane, not `undefined`,
and both call sites' `?? "todo"` fallbacks are narrower than they look.
Two of my first test cases asserted the opposite and failed; the
behaviour is the resolver's, and the write it produces is identical to
the caller's own legacy fallback either way.

## Revert proofs (measured, not asserted)

| revert | failure |
|---|---|
| `column: holdColumn` -> `column: "todo"` | `extension.test.ts`:
`expected 'todo' to be 'queued'` |
| omitted column -> `column: "triage"` | `routes-gitlab.test.ts`:
`expected 'triage' to be undefined` |

The two neighbouring `fn_delegate_task` cases stay green under the first
revert, because the built-in board and the test's `linearWorkflowIr`
both call the lane `todo` — which is exactly why this literal survived
every previous pass.

The GitLab case asserts **absence** of the key rather than a resolved
id: the store there is a fake whose `createTask` echoes its input, so
asserting a resolved value would be testing the fake. Absence is the
property that hands the decision to the real `createTask`.

## Census

| file | before | after |
|---|---|---|
| `packages/dashboard/src/routes/register-gitlab.ts` | 1 | 0 |
| `packages/cli/src/extension.ts` | 1 | 0 |

Baseline tightened. It also picks up
`packages/core/src/task-store/moves.ts` 2 -> 0, which was **already true
on main** — not from this diff.

## Noted, deliberately not changed

- `validateAssignableAgentId`'s synthetic probe a few lines above still
uses `{ id: "<new>", column: "todo" }`. It feeds `isImplementationTask`,
whose `IMPLEMENTATION_TASK_COLUMNS` set an earlier worker documented as
deliberately-not-converted (converting it makes the routing policy async
— an agent-admission behaviour change). On a renamed board the probe is
now *stricter* than the real destination, which is the safe direction
and matches the pre-existing behaviour.
- The third site from this bucket, `workflow-node-handlers.ts:455`
(`transitionTask({ column: "in-review" })` on the `review-handoff`
seam), is a **hard** failure rather than a silent one — `transitionTask`
routes through `moveTask`, which post-U12 throws
`TransitionRejectionError` for an undeclared destination, so the
workflow walk dies at the handoff on any renamed review lane. It is
engine (`batch-engine`) and fixing it properly touches `executor.ts`,
which #2820 is also editing. Left for that batch rather than opened as a
conflicting edit.

## Verification

- `pnpm test:gate` — 161 / 487 / 13 / 71 passed
- `pnpm lint` — clean
- `tsc --noEmit -p tsconfig.json` for `@fusion/core`,
`@runfusion/fusion`, `@fusion/dashboard` — clean
- `node scripts/lifecycle-column-census.mjs --strict` — exit 0
- targeted: `project-lane-vocabulary.test.ts` 14/14,
`routes-gitlab.test.ts` 8/8, `extension.test.ts -t fn_delegate_task` 9/9

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* GitLab-imported cards now appear in the workflow’s configured intake
lane.
* Delegated tasks now move to the workflow’s configured hold lane,
including workflows with custom lane names or separate intake and hold
lanes.
* Delegation reports the task’s final lane and provides an error when it
cannot be moved successfully.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 16:20:38 -07:00
gsxdsm
ea477f3ada fix(core): team analytics reported an idle project on a renamed board (#2864)
First of the 14 genuinely lane-bound SQL sites from #2839. I had been
deferring these as "the owner is active in those files" — then checked,
and **no open PR touches them**. The batch-core commit had already
landed, so there was nothing in flight to collide with. The deferral was
an assumption I could have tested two turns earlier.

## The defect

`aggregateTeamAnalytics` filtered in SQL on `"column" = 'done'` and `IN
('in-progress','in-review')`. On a board whose lanes are renamed those
match nothing, so per-agent completed counts, the project total, and the
in-flight breakdown all came back **zero**.

Nothing errors. A dashboard reading *"0 tasks completed"* for a team
that shipped all week looks like an idle project, not a bug — which is
why this survives review and never gets filed.

**Why the sweep missed it:** the lifecycle census parses TypeScript
comparisons; these ids live inside SQL strings, which are string data.
The batch-core conversion fixed this file's TS guards **today** and left
the queries untouched. The file scored as converted.

## Resolved per project, not per task

`resolveProjectColumnsForRoles` gives the union of a role's columns
across the project's workflows, which is the right set here because
analytics aggregates a whole project — so a bound `IN` list is
sufficient. The merge-queue cleanup needed the
*superset-then-decide-in-JS* shape instead, because its lanes are
genuinely per task and SQL cannot know a task's workflow. Same program,
two correct answers; worth not copying the wrong one.

## Measured

Reverted, exactly one case flips:

```
✓ default vocabulary: a completed task is counted
× renamed vocabulary: a task in the RENAMED complete lane is counted
✓ renamed vocabulary: a task in the WIP lane is NOT counted as completed
✓ without a lane store, the legacy ids still answer
  Tests  1 failed | 3 passed (4)
```

The three controls are deliberate: the default vocabulary (a generally
broken aggregator cannot hide behind the renamed case), a WIP task that
must **not** count as completed on the renamed board (resolving real
lanes must not degrade into "every column counts" — an undercount turned
overcount is harder to notice), and an omitted lane store that must keep
the legacy answer byte-identical.

## Two mistakes the first attempt made, both caught by running it

- **`= ANY(${array})` does not work.** Drizzle expands a JS array in a
template into a comma tuple, so PostgreSQL rejected `(($1,$2,$3))` with
*"op ANY/ALL (array) requires array on right side"*. An `IN` list of
individual bound parameters is the working shape. Each id stays a
parameter — these come from operator-authored workflow definitions and
are never interpolated as SQL text.
- The fixture's `ON CONFLICT (id)` had no matching constraint on
`project.agents`; the sibling suite seeds with explicit
`created_at`/`updated_at`.

## Scope

One production caller (`register-command-center-routes.ts`), threaded in
this same change so the parameter has a supplier from the start rather
than becoming another inert seam of exactly the class this program keeps
finding.

The sync SQLite arm in the same file keeps its literals: that path
throws in backend mode and has no production caller, the same dead-arm
conclusion reached for `cleanupStaleMergeQueueRowsImpl` on #2839.

## Verification

`pnpm test:gate` green · both Command Center analytics suites 8/8 ·
`tsc` core 0, dashboard 0 · lint 0 · changeset included.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 16:09:14 -07:00
gsxdsm
9e1deffa72 fix(reliability): the review-failure headline read 0% because both inputs were zero (#2861)
The Reliability panel's headline metric is computed from two queries
that name lanes:

```ts
scopedStore.getTaskMovedCountsByDay({ …, toColumn: "in-review" }),
scopedStore.getTaskMovedCountsByDay({ …, fromColumn: "in-review", toColumn: "in-progress" }),
…
const headline = inReviewFailureRate7d(enteredByDay, bouncedByDay, nowMs);
```

On a board that renamed either lane, **both return `{}`**. Every per-day
count is zero, and the headline then divides one zero by another and
reports a healthy rate.

**This is the worst shape a lifecycle defect takes.** It produces a
*number*, not an error, and the number is reassuring. An operator
reading a 0% review-failure rate beside a populated audit event list has
no reason to suspect the metric is blind — the same failure mode as the
analytics `tasksInProgress: 0` sitting next to correct cost totals.

## Why a union is correct here, not a compromise

`getTaskMovedCountsByDay` takes **one** column per side, so the lanes
are resolved to sets and the query is issued per `(from, to)` pair and
summed.

The important part is *why* summing over a union is the right answer
rather than a widening hack. These read **move history**, and a past
move recorded the column name as it was at the time — the same reasoning
that keeps `tasksEnteredInReviewPerDay` in this module matching recorded
values verbatim, marked DELIBERATE-LITERAL. A board renamed last month
therefore has old rows under the old id and new rows under the new one,
so the correct query covers **both**. That is exactly the set
`resolveProjectColumnsForRoles` returns, with the legacy id always
unioned in.

Asking for either name alone is the bug — not a choice between them.

No double-counting: a move event has exactly one `(from, to)` pair, so
the queries partition the events rather than overlapping.

**The common path does not get more expensive.** On the built-in board
this issues the same two queries as before, and there is a test
asserting the call count so a future change cannot quietly turn one
query into N.

## Revert proof (measured)

Collapse the sets back to the single literals:

```
AssertionError: expected { '2026-07-01': 1 } to deeply equal { '2026-07-01': 2, '2026-07-02': 3 }
AssertionError: expected { '2026-07-01': 1 } to deeply equal { '2026-07-01': 2, '2026-07-03': 5 }
```

The helpers live in `reliability-metrics.ts` rather than inline in
`server.ts` specifically so they are testable without booting an express
app.

## Verification

- `pnpm test:gate` — 161 / 487 / 13 / 71 passed
- `pnpm smoke:boot` — PASS (`fn --help`, real `serve` `/api/health` 200,
clean shutdown)
- `pnpm lint` — clean
- `tsc --noEmit` (`@fusion/dashboard`) — clean
- `reliability-metrics.test.ts` — 15 passed
- census `--strict` — exit 0 (unchanged: query filters are not
comparisons)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 15:57:51 -07:00
gsxdsm
eed8ca55fc batch-dashboard-src: the planner metrics tool froze active runtime on a renamed execution lane (186 → 185) (#2842)
`packages/cli` and the plugin packages are at **zero** lifecycle guards,
so this picks up the nearest unowned work: the `packages/dashboard/src/`
remainder.

## The defect

`activeRuntimeMs` adds the wall-clock since `executionStartedAt` only
while the card is accruing work — the **WIP role** — but it was keyed on
the literal `in-progress`. On a board whose execution lane is renamed,
that live tail was dropped, so `fn_task_planner_get_task_metrics`
reported active time frozen at whatever the last completed segment left
in `cumulativeActiveMs`. The number stayed plausible, which is why
nothing surfaced it.

## The part worth reading: the wiring had no watcher, from either
direction

I wired the producer (`chat.ts` resolves the task's own lanes via
`wipColumnsForTask`) in the same commit, then checked whether that
wiring was actually covered. It was not:

- **Deleting the `wipColumns:` argument left the entire 3830-test
dashboard suite green.** The formatter's own tests inject the set by
hand, so they prove the *guard* and are structurally blind to whether
production fills it.
- **`check-inert-flag-seams.mjs` does not see it either.** It tracks
trailing optional **parameters**; this is a property inside an options
bag. That is a real gap in the checker — every seam expressed as an
options-bag property is currently unguarded. Reported here rather than
fixed, because #2822 and #2830 both already modify that script and a
third change would guarantee a three-way conflict.

So `createTaskPlannerMetricsTool` is exported and a second test drives
it, letting it do its **own** resolution against a renamed board.
Deleting the argument now fails 1 of 2.

## Census

| | before | after |
|---|---|---|
| COLUMN guards | 186 | **185** |
| `packages/dashboard/src/task-planner-chat-metrics.ts` | 1 | **0** |

Baseline re-recorded; `--strict` exits 0.

## Two findings I did NOT act on, deliberately

**1. `github-tracking-state.ts` keeps 2 counted guards and should.**
They are the documented degraded-mode arms of a fully-resolved
classifier (`completeLanes === undefined ? columnId === "done" : ...`).
Marking them `DELIBERATE-LITERAL` would drop the count by
**reclassification rather than conversion** — the exact move the
census's own strict-check warns about. Related: the census reports `0
are trait-fallback branches (already converted)`, yet these are
precisely that shape, so the trait-fallback classifier appears not to
recognise a ternary whose fallback arm is the literal. Worth a look by
whoever owns the census.

**2. Three pre-existing failures in `packages/dashboard/src/__tests__`,
unrelated to this change** — measured identically on `origin/main`
before and after:
- `planning-browser-e2e.test.ts:353`
- `register-model-routes-kimi-k3-supplemental.test.ts:60`
- `routes-tasks-near-duplicate.test.ts:274`

Flagging rather than touching them; per the standing rule they are
quarantine candidates, not appeasement candidates.

## Verification

Dashboard `tsc` clean, `pnpm lint` clean, census `--strict` 0,
`check-inert-flag-seams` 21/21 supplied, changeset lint clean. Targeted
suites: `task-planner-chat-metrics.test.ts` 8/8,
`task-planner-metrics-tool-wip-lanes.test.ts` 2/2.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 15:16:24 -07:00
gsxdsm
51934931e1 fix(board): the awaitingPlanning badge only ever worked on a lane named "todo" (#2845)
Converts the one site in `register-task-workflow-routes.ts` that a
previous pass **deliberately deferred**, and does it in the shape that
note asked for.

## What the deferral said

```
FNXC:WorkflowResolvedColumns 2026-07-29-00:00 (U12 — R8, DELIBERATELY NOT CONVERTED):
… I converted it and then REVERTED: resolving each task's hold column needs a
per-task workflow read, and this is the board-load path whose own comment above
exists because unbounded reads here "turn a board load into thousands of reads".

Converting it properly needs the hold column resolved per WORKFLOW from data the
board payload already carries, not per task from the store.
```

That was the right call and the right diagnosis.
`resolveProjectColumnsForRoles(store, ["hold"])` is exactly the
project-scoped shape it names: **one** `listWorkflowDefinitions()` read
per board load, flat in task count. The expensive part — a PROMPT.md
read per row — is untouched and still bounded by
`AWAITING_PLANNING_ENRICH_LIMIT`.

The test asserts the flatness directly (`listWorkflowDefinitions` called
exactly once), so a future per-task regression fails here rather than
being discovered as board latency.

## What was broken

The filter named `todo`, so on a board whose waiting lane is called
anything else **no row was enriched at all** — no error, no log line,
just a silent fall back to the client's `steps.length === 0` heuristic.
That heuristic is precisely what this enrichment was added to correct,
so the card most likely to be mislabelled — real spec, zero parsed
steps, already a scheduler dispatch candidate — sat on "Queued to plan"
indefinitely.

Over-inclusion is the safe direction and is chosen deliberately: a card
in some other workflow's hold lane gets annotated as waiting, which is
what a waiting card in a waiting lane should show.

## Revert proof (measured)

Restore `task.column === "todo"`:

```
FAIL  register-task-workflow-routes.awaiting-planning.test.ts
  > enriches a card in a RENAMED hold lane, not only one literally named todo
  expected undefined to be false
```

The other 8 cases in the file stay green — their harness store declares
no `listWorkflowDefinitions`, so they run the degraded legacy-`todo`
path. That compatibility is half the contract, which is why the new case
brings its own store rather than widening the shared harness.

## Census

| file | before | after |
|---|---|---|
| `packages/dashboard/src/routes/register-task-workflow-routes.ts` | 3 |
2 |

The 2 remaining in that file are documented trait-fallback branches, not
unconverted debt. The baseline also picks up
`packages/core/src/task-store/moves.ts` 2 -> 0, already true on main and
not from this diff.

## Verification

- `pnpm test:gate` — 161 / 487 / 13 / 71 passed
- `pnpm lint` — clean
- `tsc --noEmit -p tsconfig.json` (`@fusion/dashboard`) — clean
- `node scripts/lifecycle-column-census.mjs --strict` — exit 0
- targeted: `register-task-workflow-routes.awaiting-planning.test.ts`
9/9

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 14:48:00 -07:00
gsxdsm
d56c635c77 test(dashboard): cover the three lane resolvers nobody was testing (#2826)
## Why this exists

#2821's review found a bug that lived entirely in a lane **builder**
while every test drove the **guard** that consumed it. That is a
structural blind spot, not a one-off: injecting a resolved value into a
synchronous guard makes the guard testable and the resolver invisible.

So I audited every lane helper I added this session. Three had **no
direct coverage at all** — `archivedColumnsForTask`,
`wipColumnsForTask`, `preWipColumnsForTask`. Their callers were tested;
the functions were not.

## They shared the defect that review named

Each read `resolved.length > 0 ? resolved : legacyId`, which conflates
two different boards:

- a **v1 upgrade** — `synthesizeDefaultColumns` emits `traits: []` on
every column, so the legacy id is the only vocabulary that exists, and
falling back is correct;
- a **v2 board that expresses traits** and declares no lane of that role
— where the legacy id names a column the board may still *have* and
deliberately did not give the role. Falling back there widens the guard
onto a role the board explicitly withheld.

`declaresAnyLifecycleTrait` separates them, matching the shape #2821's
review established for `resolveNodeOverrideLanes`.

## The fixture trap, which is the part worth reading

**My first fixture could not see the bug.** It traited the role under
test — and where the role *is* traited, the two shapes agree: both
return the traited lane. Mutating a helper back to the old shape left
all 15 cases green.

The shapes diverge only when the resolved set is **empty while traits
are expressed**. Each helper now has that case explicitly, with a
fixture that traits something *other* than the role under test.

**Mutation-verified per helper:** all three reverted independently now
fail. Before the extra case, none did.

This is the second time this session a fixture built with the production
path normalised away the very thing under test. Worth stating as a rule:
a renamed-lane fixture proves the resolver reads traits; only a
*traits-expressed-but-role-absent* fixture proves what it does when the
answer is legitimately nothing.

## Verification

- `task-lifecycle-lanes.test.ts` → 18 passed (was 15, none covering
these three)
- consumer suites (`github-issue-comment`, `planning-board-tools`,
`register-git-github.review-lanes`) → 64 passed together
- `pnpm test:gate` → 161 + 487 + 13 + 71
- `--strict` → 0; `tsc --noEmit` and `pnpm lint` → 0 errors

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 13:18:08 -07:00
gsxdsm
74cba4b46d batch-core: one shared landed-lane helper for the source-issue surfaces (75 → 72) (#2783)
## batch-core continued — the source-issue cluster

Follow-on to #2780 (merged). Scope is still `packages/core` +
`packages/dashboard/src`.

### The defect

Five places asked the same question — *has this task landed?* — and all
five compared against the literal `done`:

| surface | consequence on a renamed board |
|---|---|
| GitHub source-issue commenter | never comments on or closes the source
issue |
| GitLab source-issue commenter | same |
| GitLab `closedAt` backfill reconciler | finds nothing, reports a clean
scan |
| session-diff boundary | finished tasks diff against an already-merged
branch |
| tracking-comment transition | (already converted; left alone) |

The commenters are the sharpest case: they returned **before reading a
single setting**, so on a renamed board the feature looked *disabled*
rather than broken — an operator checking `githubCommentOnDone` would
see it enabled and still get nothing.

The backfill is the quietest: `scanned: N, filled: 0` reads as "nothing
to do", so the failure was indistinguishable from success.

### The fix

One home: `packages/dashboard/src/task-lifecycle-lanes.ts`. Callers now
only ask.

Five copies of one question is exactly how the halves drift apart — the
motivating incident is FN-6115 → FN-6118 → FN-6123, where the same
affordance was fixed three times because it lived in two components.
This also folds in the duplicate landed-lane helper I had left in
`register-session-diff-routes.ts` in the previous PR, which was the
sixth copy waiting to happen.

Two helpers, and the difference is deliberate:

- **`landedColumnsForTask`** — `complete ∪ archived`. Membership, since
a board may declare more than one column carrying either role, and
`columnsWithFlag(...)[0]` would silently ignore the second.
- **`completeColumnsForTask`** — complete only. The GitLab backfill's
own FNXC note records that archived tasks live in `archiveDb` and are
*intentionally* excluded, so it must not widen to the archived role just
because the shared helper offers it. Today it lists with
`includeArchived: false` and would see no archived rows either way — but
that is an incidental property of the query, not the contract. The test
pins the difference so the two are not later "simplified" into one,
which would change that caller's behaviour without touching it.

Both treat an **empty** resolved set as *unexpressed*, not absent — the
v1 hazard: `synthesizeDefaultColumns` upgrades a v1 graph with `traits:
[]` on every column, so reading empty as "no complete lane" would stop
these surfaces firing on every pre-v2 project.

The reconciler is two-stage on purpose: the cheap provider and
`closedAt` tests run first and reject almost everything, so a workflow
read only happens for real candidates, and it shares one IR cache across
the scan — one read per distinct workflow rather than per task.

### Census

`batch-core` scope **75 → 72**; repo total **338**.

### Verification

- `pnpm --filter @fusion/dashboard exec tsc --noEmit -p tsconfig.json` →
0 errors
- `pnpm lint` → 0 errors
- commenter + reconciler suites → **63 passed**; helper suite → **5
passed**
- **Mutation-verified:** making the helper ignore its resolved set fails
1 of 5.

---

## Round 2 — server.ts, chat.ts, and a correction

**Census: 75 → 67** across this PR.

### The correction (see the review thread above)

My first pass gated the source-issue commenters on
`landedColumnsForTask` (`complete ∪ archived`), which **widened** the
trigger — `to === "done"` never fired on archival, and the landed set
does. Both commenters now use `completeColumnsForTask`, and the unused
`hasTaskLanded` wrapper is gone.

The ratchet for it is pinned on the **default** board, deliberately: a
widening is visible exactly where the legacy names still apply, so no
renamed-board fixture would catch it.

### `chat.ts` — three sites, and a pair that had to move together

- **Chat verification** required `column === "in-progress"`, so on a
renamed board every chat-driven verification was refused with a message
naming a column the board does not have.
- **The planner refinement pair.** Two separate guards decide this
feature: `createSession` *registers* the tool only for a finished task,
and the tool's own `execute()` *refuses* a non-finished source. Both
compared `done`. Converting only one half would have offered the tool
and then had it refuse itself — the half-converted-pair shape. The new
test asserts **both** halves in one case (tool present *and* refinement
created), and each half reverted independently fails it.

Existing `chat-manager` coverage caught neither revert, which is why the
case exists rather than relying on the suite that was already there.

Complete-only again, not the landed set: an archived task is off the
board and is not a refinement source.

### `server.ts`

- **Planner-chat retention** — the archival cutoff was a literal, so on
a renamed board task-planner chat sessions were retained forever; the
rule this listener exists to enforce never fired. Resolved, and awaited
inside the existing fire-and-forget chain rather than by making the
listener `async` — `task:moved` has synchronous subscribers whose
ordering is load-bearing elsewhere, and a chat-row delete is not the
right place to introduce a microtask boundary into that emit.

- **`isBadgeEligibleTask` — deliberately NOT converted, and marked as
backlog.** On a renamed board it is genuinely wrong: an archived card
stays badge-eligible, its snapshot is never evicted, and the cache grows
for the daemon's lifetime — the exact memory leak the predicate was
added to fix, back under a different column name.

What blocks it is measured, not assumed: both callers are synchronous
`task:updated` / `task:created` listeners whose next statement is
documented as *"Update local cache immediately"*, so awaiting lets a
second event for the same task interleave between the eligibility check
and the cache write.

I did **not** add an optional `archivedColumns` parameter, because
nothing could fill it — the callers are the sync listeners. That is the
inert-injection shape this PR's own review caught twice on #2780: the
predicate would read as converted, its test would pass by injecting the
value, and production would keep the literal. The unblocking change (a
resolved-archived-lane cache on the badge-snapshot scope, keeping the
predicate synchronous) is recorded at the site.

### Verification

- `tsc --noEmit` → 0 errors; `pnpm lint` → 0 errors
- `chat-manager` → 101 passed; commenter/reconciler/helper/badge suites
→ 55 passed
- Mutation-verified per fix, including each half of the refinement pair
separately

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Improved task lifecycle handling for renamed workflow lanes, including
completed, archived, landed, and in-progress states.
* Task lists now exclude completed tasks regardless of the completion
lane’s name.
* Chat verification and refinement actions now recognize configured
workflow lanes.
* GitHub and GitLab completion comments trigger only for genuinely
completed tasks, not archived tasks.
* Knowledge index refreshes and GitLab metadata updates now support
custom completion lanes.
* **Tests**
* Added regression coverage for renamed completion lanes and
archived-task behavior.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 12:03:07 -07:00
gsxdsm
e84e9d7f60 fix: the caller audit — five unwired parameters, five defects in their callers (#2803)
Seven fixes that were sitting on separate handoff branches with no owner
while `main` moved. Consolidated, rebased onto current `main`, and
verified **together** rather than only per-branch. The individual
branches remain if a subset is preferred.

This is the same consolidation that got `batch-core` and #2787 adopted.
**Close it if it breaks queue policy** — the branch keeps the work safe
either way.

## Where these came from

#2787's review found an optional parameter whose production caller never
passed it. That is a class, so I ran it against everything I had landed
and found five more. **All five turned out to have their real defect in
the CALLER, not the parameter** — in four of them the parameter was
unreachable:

| unwired parameter | what was actually wrong |
|---|---|
| `blocker-fanout.escalationColumns` | the hold default made the count
zero — **no bottleneck warning was emitted at all** |
| analytics `columnFlagsByName` | routes never built a map — **0
in-progress / 0 in-review beside correct cost totals** |
| `isLegacyAutoMergeStampCandidate` | the read **queried a column a
renamed board does not have**, so the backfill iterated nothing |
| `rankAssignedTasksForWakeDelta` | `getTasksByAssignedAgent`'s
`excludeArchived` used the literal — **archived cards returned as open
work** |
| `duplicate-intake.columnFlagsByColumnId` | intake could **archive or
soft-delete a newly created task** as a duplicate of finished work |

The heuristic worth keeping: **an optional parameter no production
caller fills is a marker pointing at an unexamined caller.** The census
cannot see any of these five — every gate is a `Set`/array literal or a
query filter, i.e. a definition rather than a comparison.

## Also included

- **`executor.ts`** — the stale-spec guard did the exact thing its own
comment forbids: on a renamed board it ran on a LIVE task and pulled it
out of execution into replan. `activeMergeStatuses` protected merging
cards *by accident*, which is why the symptom looked arbitrary.
- **`register-project-routes.ts`** — project health reported **0 active
tasks**; its list also still contained `triage`, dead since U11.
- **`dashboard/app/utils/taskTiming.ts`** — a **second copy** of
`getTotalAgentActiveMs`. Core's was converted; the card chip imports
this one, so the census counted the site as done while the rendered
number stayed keyed on `"in-progress"`.

## Verification

Verified as a set: `pnpm test:gate` **161 / 13 / 487 / 71** · core
suites **15 passed** · engine **7** · dashboard **12** · four `tsc`
targets clean · lint clean · census `--strict` exits 0.

Each fix is revert-proven individually; the specific case that fails is
named in each test header.

## Two honesty notes

**Three guards here are structural, not behavioural, and say so in their
headers.** `sanitizeAgentTaskLinks` is a closure inside
`createApiRoutes`; the analytics aggregators need a live
`AsyncDataLayer`; the stale-spec guard sits deep inside `execute()`.
Each ratchet fails on revert — verified — but none is an end-to-end
proof, and the headers state which half they cover.

**One of my behavioural test sets would have lied.** The intake-dedup
cases drive `findSameAgentDuplicates` directly; I removed the wiring to
measure the revert and **they stayed green**, because they pin the
predicate and not the caller. That is the exact illusion this audit was
chasing, reproduced in my own file. The forward now has its own
structural check.

## Deliberately not included

`worktree-pool.ts:1205` — it **fails safe** (a missed match protects a
branch from cleanup rather than deleting it) and sits in the merger's
branch-reaping path where the opposite error destroys work. That
deserves its owner's judgement, not a drive-by conversion.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 12:02:53 -07:00
gsxdsm
6bdde6f246 fix: five lifecycle gates the census cannot see — incl. live ephemeral workers reaped and duplicate follow-up cards (#2787)
Five lifecycle-column fixes the census **structurally cannot see**. Each
gate is a `Set` or array literal — a *definition*, not a comparison — so
no backlog entry ever pointed at any of these files. Found by grepping
for lane-shaped list literals after the same shape surfaced in
`duplicate-intake` and `blocker-fanout` (both merged via #2780), then
confirmed by reading each USE site.

**On opening this:** I offered twice to fold these into a PR and kept
them on handoff refs to respect one-open-PR-per-worker. They have now
sat unadopted across several cycles while `main` moved, and two of them
destroy or duplicate work. Opening is the reversible call — **close it
if it breaks queue policy** and I will keep them on the branch.

## What is in it

| commit | defect on a renamed board | severity |
|---|---|---|
| `beb107a7bc` | assignment load-balancing **defeated** —
`assignmentLoad` stays empty, every candidate reads as load 0, the sort
falls through to its stable `createdAt` tiebreak, so **one agent wins
every assignment** while the rest idle | distribution |
| `cf4b59e1cb` | the zombie sweep **deletes LIVE ephemeral workers** |
**destroys work** |
| `5fe004ae64` | eval follow-up dedup sees **zero open tasks**, so every
run re-files follow-ups it already filed | **duplicate cards** |
| `a1021de8b2` | agents keep a **"working on" indicator for finished
cards** | stale UI |
| `86680d1220` | the **Files tab never loads** — the fetch never fires |
silent empty |

### The one that destroys work

`shouldDeleteOnSweep` tested a hard-coded terminal `Set`, then fell
through to `return task.column !== "in-progress"`. On a renamed board
**both halves miss, and they compound in the worst order**: the terminal
test fails, control reaches the fallthrough, and `"building" !==
"in-progress"` is `true`. An ephemeral worker **actively executing a
task** is classified as a zombie and deleted. Nothing logs.

Its fallback is **deliberately asymmetric**, and the comment says why:
an unresolvable workflow keeps the legacy literals rather than guessing.
Failing to reap a dead worker costs a slot; reaping a live one destroys
work in flight. Those are not symmetric, so uncertainty fails toward
keeping the worker.

## Verification

Verified **as a set**, not only per-branch:

- `pnpm test:gate` — **161 / 13 / 487 / 71**
- engine suites (assignment, ephemeral, eval-followups) — **44 passed**
- dashboard suites (agent-task-link, useSessionFiles) — **16 passed**
- `tsc` engine + dashboard server + dashboard app — clean
- `pnpm lint` clean · census `--strict` exits 0

**Revert-proven individually.** Restoring each literal fails its own
case: the renamed-wip zombie case, the renamed-wip assignment case, the
renamed-lane dedup case, the sanitizer ratchet, and both
`useSessionFiles` role cases.

## Two honesty notes, flagged rather than buried

**`a1021de8b2`'s guard is STRUCTURAL, not behavioural.**
`sanitizeAgentTaskLinks` is a closure inside `createApiRoutes`,
reachable only by standing up the full express app. The ratchet asserts
the source — resolver threaded per task, bare literal call gone, cache
shared, fallback retained — and **fails on revert**, verified. It is not
a substitute for a behavioural test; whoever owns the dashboard server
should add one if that seam grows.

**`useSessionFiles`'s negative case passed in isolation and failed in
the suite.** Hooks are not unmounted between cases there, so a prior
case's in-flight fetch landed inside it. That is the classic shape of a
test that gets "fixed" by reordering; it now asserts a **delta** against
the pre-render call count, which is independent of what leaks in.

## Deliberately NOT included

`worktree-pool.ts:1205` — the sixth site from the same sweep. It **fails
safe**: a missed match means the skip does not fire, so the branch is
added to `activeBranches` and *protected* from cleanup. The cost is
stale branches accumulating, not deletion. It also sits in the merger's
branch-reaping path, where the opposite error destroys work, so it
deserves its owner's judgement rather than a drive-by conversion.
Flagged, not guessed.

Also still open and unclaimed: roughly 69 untriaged literal-list sites
across engine/dashboard/cli. The grep is one line and the file list is
on #2775 — with the measured caveat that about half are false positives
on shape alone (`LEGACY_*` names, seeds unioned with resolved values,
and `roles: ["triage"]`, which is an `AgentCapability`, not the deleted
column). Only the use site settles it.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 10:35:31 -07:00
gsxdsm
8c9b84ae38 batch-core: packages/core + dashboard/src lifecycle conversion (129 → 92) (#2780)
## batch-core — `packages/core` + `packages/dashboard/src`

Shared branch: two workers are converting into it. Opening the PR
because the branch was green with none, and a branch without a PR merges
nothing.

### Census

Measured with `node scripts/lifecycle-column-census.mjs --json`.

| | guards |
|---|---|
| batch-core scope at branch point | 129 |
| batch-core scope now | **92** (51 files) |
| repo total now | 358 |

Files closed so far: `store.ts` 11→0, `task-merge.ts` 6→0,
`live-agent-count.ts` 6→0 (marked, not converted — see #2762),
`task-update.ts` 3→0, display-ordering + Wake Delta ranking 5→0,
`register-git-github.ts` 4→0.

### The `register-git-github.ts` slice

Three PR routes — `pr/create`, `pr/push-branch`, `pr/resolve-conflicts`
— plus the `CHANGES_REQUESTED` handler each compared `task.column !==
"in-review"`. On a renamed board **none** of them matched, so every PR
affordance the dashboard offers was refused for a card sitting in the
lane that board calls review, and the refusal named a column that does
not exist there.

All four now share one helper, `reviewColumnsForTask`, which gets two
things right that this program has repeatedly gotten wrong:

- **Membership, not a single id.** It takes the broad review set
(`mergeOrchestration ∪ mergeBlocker ∪ humanReview`).
`resolveLifecycleColumns` returns the *first* column per trait, so a
single-id answer silently ignores a board that declares a merge lane
**and** a separate human sign-off lane. These guards only refuse or
permit — they never move the card — so over-admitting costs nothing
while under-admitting refuses a request that should have worked.
- **An empty resolved set means UNEXPRESSED, not absent.**
`synthesizeDefaultColumns` upgrades a v1 graph by emitting every default
column with `traits: []`, so a v1-upgraded workflow resolves to an empty
review set while its `in-review` column plainly exists and holds the
card. Reading empty as "this board has no review lane" would refuse
these routes on **every pre-v2 project** — a worse regression than the
one being fixed, and invisible to any v2 test.

This is the dashboard twin of the `fn pr create` guard in
`packages/cli/src/commands/pr.ts` (#2775). The two surfaces answer the
same question and now agree — FN-5893 surface enumeration.

### Testing note: why the seam and not the routes

I wrote route-level HTTP tests first and **deleted them**. An express
fixture over `registerGitGitHubRoutes` hangs — every case, including the
pure refusals, times out at 4s, because registering the router starts
background work the fixture never satisfies. Making it run would mean
mocking git, the GitHub client, and the pollers: a mock-the-world shell,
which is what the project's do-not-add-slow-tests rule (FN-5048) says to
avoid in favour of a narrow seam.

`reviewColumnsForTask` *is* the narrow seam — it holds the entire
decision, and the four call sites now do nothing but ask it and render
its answer. Six cases pin it: the renamed lane is returned and
`in-review` is not, a two-lane board returns both, a v1-upgraded board
falls back, an unresolvable workflow falls back, and the refusal renders
lanes an operator can act on.

**Mutation-verified, both directions:** reverting the helper to the
legacy literal fails 2 of 6; treating an empty set as an answer fails 1
of 6.

One fixture bug worth recording, since it would have made the two-lane
case vacuous: the trait id is kebab-case `human-review`, not
`humanReview`, and the built-in traits must be registered via `import
"@fusion/core"` before flags resolve.

### Verification

- `pnpm --filter @fusion/dashboard exec tsc --noEmit -p tsconfig.json` →
0 errors
- `pnpm lint` → 0 errors
- `register-git-github.review-lanes.test.ts` → 6 passed

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 09:28:27 -07:00
gsxdsm
c428eed9d7 fix(test): two dashboard-api reds — conversions changed the log CHANNEL and the error MESSAGE (#2774)
Both red on main. `api:curated` goes **2 failed → 34 files / 1599
passed**. Neither is a product defect — both are conversions the tests
had not followed.

## 1. The log channel moved

`sse.test.ts` spied `console.log`. `sseDebug` routes through
`createLogger("sse").debug` (`sse.ts:50-53`), and the shared logger
writes debug lines to **`console.error`** carrying a `\0fnlvl=info\0`
severity marker — that is the point of FN-8603's adapter.

So the spy saw nothing, and the failure read `expected false to be
true`, naming neither the channel nor the logger. The stderr in the run
output showed the lines being emitted the whole time:

```
fnlvl=info [sse] [sse] + connection (active=1, hwm=2)
fnlvl=info [sse] [sse] - connection (active=0)
```

## 2. The error message is now built from resolved lanes

`routes-tasks` asserted the substring `"in-review or in-progress"`. The
message is now:

```ts
const allowed = [...prFeedbackReviewColumns, prFeedbackWipColumn]
  .map((column) => `'${column}'`).join(" or ");
throw badRequest(`PR feedback can only be addressed for tasks in ${allowed}`);
```

so it reads `'in-review' or 'in-progress'` — quoted, and derived from
the resolved columns.

**Asserted each lane separately rather than re-pinning the joined
string.** The join order and separator are presentation; the lanes being
the resolved review + wip columns is the fact this case owns. Re-pinning
the punctuation would break again on the next formatting change *and*
would not have caught a wrong lane — which is the failure this test
exists to catch on a renamed board.

## Verification

| check | result |
|---|---|
| `test:quality:api:curated` | 2 failed → **34 files / 1599 passed** |
| `sse.test.ts` | **24 passed** |
| `routes-tasks.test.ts` | **99 passed** |
| `pnpm lint`, dashboard `tsc` | clean |

## Scope

Fix-forward only, per the u9 lane. Found by re-scanning the packages
after #2739 / #2744 / #2754 merged, rather than by waiting for a report.

For the record on the other groups at the same commit: `components-a`
**1195 passed**, core is **2 failed** — both already accounted for
(`archived-column-gate-parity` is #2768's target,
`agent-logs-and-monitor.pg` is the deferred funnel/analytics decision on
#2669).
2026-07-30 09:22:05 -07:00
gsxdsm
2e4905fa0e refactor: one definition of "which columns are review" — three copies deleted onto core's resolver (#2751)
**#2730 added `resolveReviewColumns` to core. This deletes the three
copies that predated it.**

Measured on `origin/main` before this change — three in-tree
definitions, **none of which agreed**:

| site | definition |
|---|---|
| `core/workflow-lifecycle-traits.ts` (#2730, authoritative) |
mergeOrchestration ∪ mergeBlocker ∪ humanReview — **all** columns |
| `dashboard/routes/register-task-workflow-routes.ts` | mergeBlocker ∪
humanReview ∪ **first** mergeOrchestration |
| `cli/src/extension.ts` | mergeBlocker ∪ humanReview ∪ **first**
mergeOrchestration |
| `cli/src/commands/task.ts` | all three, full union |

**Both `.slice(0, 1)` variants are mine**, from #2723's review round: I
narrowed to core's then-single `.review` because the reviewer was right
that a superset let the dashboard act on a lane the engine did not own.
#2730 answered that question authoritatively in the other direction, so
the narrowing is obsolete.

Worse, and the part that makes this urgent rather than tidy: **the two
CLI copies had already drifted apart inside #2728.** `fn_task_retry`
refused a card in a second merge lane that `fn task retry` accepted —
two surfaces, one operator action, two answers, from two copies of one
definition written days apart by me.

All three now call core. The dashboard keeps its thin store→IR wrapper
(its callers hold a store and a task id, not an IR) but the **body** is
core's.

## One assertion inverted, deliberately

My #2723 case asserted that a **second** `mergeOrchestration` column is
**refused**. Core says every merge lane is review, so the behaviour
legitimately changed and the assertion flips with it.

**Kept rather than deleted**, because the invariant under test — *the
routes agree with core* — is unchanged. Deleting the case would have
hidden that its answer moved; inverting it records which decision moved
and why. A test whose expectation quietly disappears is
indistinguishable from a test that was wrong.

## A footgun found while rebasing

The shipped signature is
`isInReviewMissingWorktreeSessionStartFailure(task, isReviewColumn?:
boolean)` — the merged version takes the **answer**, not the lanes. My
branch had passed a `ReadonlySet`, and because the parameter is `boolean
| undefined` with a `??` default, **a truthy object makes it answer
`true` for every column**.

TypeScript stops typed callers; my test only reached it through an `as
never` cast, which is how I found it. All three production call sites
correctly pass `retryReviewColumns.has(task.column)` — now asserted
structurally so a fourth surface cannot omit it.

The boolean is arguably the better shape, and I'd keep it: there is
nothing left for the callee to re-derive, so it cannot disagree with the
caller's own membership test.

## The ratchet

No surface may reintroduce a local review union (`columnsWithFlag(…,
"mergeBlocker" | "humanReview")`). Those three copies appeared because
each was added **in good faith, in a different review round, by someone
reading only their own call site** — which no amount of care prevents
and a ratchet does.

## Verification

census **553** · `pnpm test:gate` **487 / 10 / 71** · `tsc` clean in cli
and dashboard · `pnpm lint` clean · 10/10 in each touched suite.

**Pre-existing, not mine:**
`register-task-workflow-routes.move-bypassguards.test.ts` fails on
`origin/main` (400 vs 200) — already reported on #2723.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 08:28:32 -07:00
gsxdsm
d86c1f9d29 batch-engine: packages/engine lifecycle-column conversions (capacity worker's mega-batch) (#2773)
The engine mega-batch. Folds my four engine PRs and will absorb the
remaining `packages/engine` guards as commits on this branch.

**Superseded and closed:** #2722, #2741, #2766, #2770.

## Census — files converted so far

| file | before | after |
|---|---:|---:|
| `notification/notification-service.ts` | 9 | **5** |
| `runtimes/in-process-runtime.ts` | 6 | **1** |
| `eval-followups.ts` | 2 | **0** |
| `pr-comment-handler.ts` | 1 | **0** |
| `task-revert.ts` | 2 | **0** |

The last two are **census-invisible** (`Set.has(task.column)`
membership) — the class measured in #2763, which a comparison-based scan
cannot count. So the backlog number moves less than the work does,
deliberately.

## What each one actually fixed — all silent, none cosmetic

- **Notifications stopped entirely.** `handleTaskMovedAsync` compared
`data.to` to `in-review`/`done`, so on a renamed board the two
notifications operators rely on most were never sent.
- **A finished card's plan review could re-enter.** The continuation
drain's terminal test matched nothing, so a completed card's planning
continuation was handed to the executor.
- **The revert route admitted and the service refused.** The route
resolved terminal lanes; the service compared to a hardcoded pair. The
operator got a dead end from an affordance the UI and route both
offered.
- **Follow-up dedup blocked new cards forever.** A finished follow-up in
a renamed complete lane read as *open*, so the dedup matched it
permanently — defeating the intent the code documents in the line above
it.
- **The mission requeue wrote a column that may not exist**, and its
guard never matched.

## Flagged, not fixed — deliberately

- **`concurrency.ts` idle semaphore leak recovery** — the last live
caller of the running-agent predicate that does not enrich. On a renamed
board it under-counts and can reclaim a legitimately-held slot. The
enriching variant is async and this is a synchronous repair path whose
failure mode is reclaiming live work.
- **The archival `task:moved` listener** — runs on every move with no
cheap gate ahead of it; converting costs an IR resolution per move to
decide most moves are not archival.

## Notes carried from the folded PRs

Two conflicts resolved in main's favour because **main's version was
better**: `in-process-runtime`'s seam uses `terminalColumns:
ReadonlySet` (membership) where mine used `LifecycleColumns`
(first-per-role), and the test is rewritten against main's API. That
arity trap has now caught me four times, so membership is the default
shape in everything new here.

Review fixes from the folded PRs are included: the notifier's review
set, the second human-review site, the second dedup copy, the workspace
revert surface, and the file-content assertions.

## Verification

`pnpm test:gate` **GREEN** (158 + 10 + 487 + 71) · **224 passed** across
the touched engine suites · engine and dashboard `tsc` clean · `pnpm
lint` clean · census `--strict` exits 0.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 08:24:50 -07:00
gsxdsm
54d1a29621 fix(dashboard): the github-tracking-state classifier seam was never wired — a renamed terminal column still never closed its issue (#2754)
Not a conversion. A **conversion that was never connected to
production**, found by a parity sweep rather than by the census.

## How it surfaced

An AST pass over all **43** github/gitlab-named files, paired by name
with counts compared:

```
github-tracking-comments.ts=9  vs  gitlab-tracking-comments.ts=4   ASYMMETRIC
github-tracking-state.ts=2     vs  gitlab-tracking-state.ts=0      ASYMMETRIC
github-issue-comment.ts=1      vs  gitlab-issue-comment.ts=1
...8 more pairs, all symmetric at 0
```

The first is the pair #2715 fixes. The second pointed here — and the
asymmetry turned out not to be the interesting part.

## The defect

`decideIssueAction` has accepted an injectable `classify` since U12/R2,
and that file's own header states the bug the seam fixed:

> "A user-authored workflow whose terminal column is called something
else never closed its linked GitHub issue, and a custom archive column
never mapped to `not_planned`."

Its **only production caller** passed no classifier:

```ts
const decision = decideIssueAction(event.from, event.to);
```

So every real move fell through to `legacyColumnLifecycleClass`, and
**the documented bug was still live**. The seam was reachable from unit
tests only — which is why all 68 cases in that file were green while the
behaviour they document did not work.

Same shape as this branch's earlier finding on the tracking-comment
guard, where the guard returned *before* resolving. **Adding a seam and
wiring it are two changes; only the second one fixes anything.** Worth
watching for elsewhere in this program: a file can read as fully
converted, pass its suite, and still take the legacy path on every call.

## Ordering, inverted on purpose

`decideIssueAction` ran first, before the tracking-enabled check,
because comparing two strings is free. Resolving a workflow is not — so
the cheap property read now short-circuits and only tracked tasks
resolve, the ordering `github-tracking-comments.ts` and its GitLab twin
already settled on. Untracked tasks returned without acting before and
still do.

The two remaining literals **are** `legacyColumnLifecycleClass`, that
seam's named default, now marked `DELIBERATE-LITERAL` — and marked only
in the same commit as the wiring. While the default was the live path on
every move, exempting it would have hidden the real defect behind a
marker.

## Revert proof — it detects an *unwired* seam, not a missing one

Dropping the resolved classifier while **leaving the seam intact** fails
both new cases with 0 `setIssueState` calls. That is the whole point:
the new cases drive the **service**, not the pure decision function, so
they fail for exactly the reason the 68 existing cases could not. Those
pass either way.

## Two self-inflicted errors, recorded because both are recurrences

- **I hand-edited the baseline with python and wrote a raw NUL byte into
the JSON**, breaking the census parse. The `deliberateByFile` keys use a
real `\0` separator and must be written through `JSON.stringify`, never
string interpolation.
- **I then restored the baseline from a newer `origin/main` than my
branch point**, which made `--strict` report a `scheduler.ts: 12 → 26`
rise that was pure version mixing. Rebase first, then edit. (The real
`scheduler.ts` baseline staleness is already owned by #2712 — I checked
before assuming it was mine.)

## Verification

`pnpm test:gate` **GREEN** (158 + 10 + 487 + 71) · **71 passed** in the
tracking-state suite · dashboard `tsc` clean · `pnpm lint` clean ·
census `--strict` exits 0.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 08:05:50 -07:00
gsxdsm
58f1e41aa0 test(dashboard): keep the renamed-board tracking suite — its source changes landed as #2715 and #2737 (#2714)
**Claim announced on #2706 before starting.**
`github-tracking-comments.ts` + `github-tracking-reconciler.ts` — one
coherent subsystem, 9 guards each.

**18 → 7 by census, 18 → 0 behaviourally.** The gap is explained at the
bottom and it is not hand-waving.

## Both halves failed quietly, in the way that suppresses its own
evidence

| surface | what a renamed board got |
|---|---|
| the comment poster | returned early for **every** move — the tracked
issue silently stopped receiving both its "in progress" and its "done"
comment. The operator sees a linked GitHub issue that never updates. |
| the reconciler (3 scan passes) | matched **zero** tasks, so completed
work's issues were never closed — and the pass reported a clean
`scanned: 0`. |

That second one is the shape worth internalising: **the number that
would have revealed the problem is the number the bug suppresses.** No
error, no warning, a green sweep.

## The comment poster needed a derivation, not a swap

`event.to` was **both** compared against the two literals **and** passed
into `formatTrackingComment` as its `transition` argument (typed
`"in-progress" | "done"`). One value carrying two meanings: a lane id
and a comment kind.

Eight independent swaps would have had to keep agreeing with each other
forever — and a ninth site (the template's own `transition === "done"`)
is *not* a column at all, so a mechanical sweep would have converted it
wrongly. Resolving the lanes once and deriving the kind separates the
two meanings permanently. Log details still print the real column, so
the operator reads their own board's name.

## The reconciler

Per task through **one shared IR cache per scan**, resolved into a `Set`
of terminal ids rather than an async predicate inside `.filter(...)` —
`Array.filter` ignores promises, so an async predicate there silently
keeps **every** row. That is a trap worth naming for other fleet workers
converting list filters.

The archived-vs-complete distinction keeps its own resolver rather than
reusing the terminal pair: it decides GitHub's `state_reason`, and
closing a finished issue as `not_planned` is operator-visible and wrong
— as is the reverse.

## Revert proof

**5 of 8 new cases redden.** 201/201 across the nine `github-tracking`
suites (193 were already there and still pass).

## Why the census says 7 and not 0

The reconciler goes **9 → 0**. The comment poster still reports **7**,
and every one of those is `transition === "in-progress" | "done"` — the
derived comment **kind**, not a column. There is no lane comparison left
in the file.

That is exactly the vocabulary-collision class **#2692** is fixing (it
already lists five misclassified receivers: an SSE event type, a
cache-key mode, an evidence kind, a telemetry event kind, an agent
state). **`transition` is a sixth and I have reported it there.** Until
that lands the census counts them, so I am reporting both numbers rather
than the flattering one.

I deliberately did **not** mark them `DELIBERATE-LITERAL` to move the
count: that marker means "a lifecycle literal reviewed and kept", and
these are not lifecycle literals at all. Using it as a census-silencer
would put a wrong reason in the code to make a number look better.

## Verification

`pnpm test:gate` **487 / 71** · **201/201** github-tracking suites ·
`tsc -p packages/dashboard` clean · `pnpm lint` clean · census
`--strict` exit 0 (it also tightened three entries other workers' merges
left stale — the #2679 auto-tighten working).

No changeset: `@fusion/dashboard` is private and this is internal
behaviour on renamed boards.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 07:53:45 -07:00
gsxdsm
ceca08b1c3 fleet: github-tracking-reconciler 9 → 0 — deciding the sync-filter class (prefetch a resolved map), and the reconciler closed NO issues on a renamed board (#2737)
`github-tracking-reconciler.ts` 9 → **0**, and the reference
implementation for the `.filter((task) => task.column === "<id>")` shape
I have been flagging across four files.

## I stopped waiting and decided it

I flagged this class in #2709, #2696, #2700 and #2715 as "needs one
decision" and left ~25 sites unconverted. That decision was mine to make
and I should have made it three PRs ago.

**Prefetch a resolved map, then filter synchronously.** The alternative
— async predicates — forces every caller into `for await` and turns a
list comprehension into a sequential walk. Prefetching keeps the filters
synchronous, puts the awaits in one bounded place, and lets the IR cache
do the job it was explicitly built for:

> "A self-healing pass over 400 cards spanning three workflows must read
three IRs, not 400."

The cache is **instance-scoped and shared across all four passes**, so
each distinct workflow's IR is read once for the whole run rather than
once per pass. `resolveLifecycleColumns` is pure and *not* memoized by
that cache, so this still costs one cheap struct build per task — fine
in a background reconcile, and stated rather than hidden.

No new abstraction: `resolveTaskLifecycleColumns` already takes a
caller-owned cache. The only new code is a local map builder and two
named predicates.

## What it cost before

On a board with renamed terminal lanes, **every filter here matched
nothing**. The reconciler closed **no** GitHub issues and reported
`scanned: 0` — a clean-looking pass that did nothing.

## Why this is not the split brain #2724 documents — checked, not
assumed

#2724 proves the archived gate in `packages/core` is enforced in three
encodings, so converting one alone diverges them. I checked whether that
applies here before converting:

- This file contains **zero SQL** — measured: no drizzle, no `sql`
template, no `eq`/`ne`.
- It calls `listTasks({ includeArchived: true })`, so the SQL half has
already been told to include archived rows. The filter **selects among
rows it was handed** rather than deciding liveness a second time.

**Gate versus consumer** is the distinction, and a consumer can be
converted alone.

The fourth pass needed its own check because its list comes from
`listTasksForGithubTrackingReconcile`, which *is* SQL — but that impl
filters on `deletedAt IS NOT NULL` and `githubTracking IS NOT NULL`,
**never on the column**, so there is no SQL-side encoding of this
question to diverge from.

## Why the 33 existing tests stayed green through the conversion

Their fake store has **no workflow reader**, so
`resolveTaskLifecycleColumns` catches and returns `undefined` and every
case asserts the legacy fallback — exactly what it always asserted.
**None of them could have caught this being wrong.** `workflowIr` is now
an opt-in on that fake, which is what makes the new cases real tests
rather than restatements.

| reverted | result |
|---|---|
| terminal filter back to the ids | "closes issues on a RENAMED complete
lane" fails, no `setIssueState` |
| same | renamed archived-heuristic case fails, no `setIssueState` |

## A reachability finding, recorded not acted on

In backend mode `reconcileDeletedAndArchived` returns only
**soft-deleted** rows — its own comment says the archived-tasks fallback
is a separate `AsyncArchiveLineage` subsystem, skipped there — and
`task.deletedAt` is tested *first* in the `stateReason` chain. So its
archived arm is **effectively unreachable today**. I converted it rather
than deleting it: it is the documented FN-5577 done-heuristic, and
whether that fallback should be wired here is a separate question from
what vocabulary it speaks.

## Verification

`pnpm test:gate` **GREEN** (158 + 10 + 487 + 71) · **35 passed** across
the three reconciler suites · dashboard `tsc` clean · `pnpm lint` clean
· census `--strict` exits 0.

Remaining files in this class (`branch-group-ops.ts`, `store.ts`, and
the dependency pairs) can now follow this pattern instead of waiting —
with the gate-versus-consumer check applied to each, since
`branch-group-ops.ts` sits closer to the persistence layer than this one
does.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 07:14:54 -07:00
gsxdsm
15f90706e6 fleet: reliability-metrics.ts 6 → 0 — historical log values, marked not converted (#2756)
Unclaimed file, no overlap with any open fleet PR — deliberately picked
to avoid adding conflicts to the queue.

## Census

| | before | after |
|---|---|---|
| backlog | 539 | **533** |
| reviewed (DELIBERATE-LITERAL) | 31 | 36 |
| this file | 6 | **0** |

`--strict` exit 0, baseline re-recorded in the same commit.

## Why these are marked, not converted

All six ids come from `metadataColumn(entry, "from"|"to")` — the columns
**recorded on a past move event** in the activity log, not a task's
current column.

There is no workflow to resolve them against. The event was written
under whatever the board looked like at the time, and **a column renamed
since leaves every older entry carrying the old id forever.** Converting
them to a trait read would ask *"what role does the column named X play
today?"* about a record written months ago, possibly under a different
workflow — a different question with a different answer.

The failure mode matters: a trait-converted reader on a renamed board
would **zero the series** rather than fix it, silently dropping history
out of `tasksEnteredInReviewPerDay`, `tasksBouncedToInProgressPerDay`,
and `inReviewDurationMetrics`. That is worse than the literal, which at
least keeps matching the data that exists.

**The real fix for renamed boards is at the WRITER** — emit a role
alongside the id when the move event is recorded — not at this reader.
Noted at the site so whoever does that work finds it.

## A rule this generalises to

**Any reader of activity-log or run-audit metadata is a mark, not a
convert.** The census cannot distinguish `task.column === "in-review"`
(a live question, convert it) from `metadataColumn(entry, "to") ===
"in-review"` (a historical record, match it as recorded) — both are just
literals to the AST. Other fleet workers hitting log/audit readers
should expect the same call.

## Placement trap, third occurrence

My first pass marked the `const from`/`const to` declarations and moved
the count by **1 of 6** — the census excuses the construct a marker is
attached to, and the guards live in **sibling `if` statements**. Moved
the markers to the enclosing functions.

This has now caught #2645's author, me on `TaskContextMenu`, and me
again here. **Verify a marker by the count moving, not by the comment
existing** — and until every worker does, a batch reporting "N → 0" can
be off by most of N.

## Verification

Dashboard typecheck clean · reliability suites green (11 passed) ·
`--strict` exit 0 · no behavior change (comments only).

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 07:01:02 -07:00
gsxdsm
edab088107 fix(test): 23 reds on main from the column U11 deleted — routes-github (22) + move-bypassguards (1) (#2758)
**Red suite on `origin/main`, unclaimed. Reported twice (on #2723 and
#2733) and nobody picked it up, so I am fixing it rather than reporting
a third time.**

`register-task-workflow-routes.move-bypassguards.test.ts` fails on main:
`expected 400 to be 200`.

## Not a route bug — a fixture that outlived its column

The move endpoint validates the target against the **task's own
workflow** (U12/R2), and the default lineage post-#2515 declares `todo |
in-progress | in-review | done | archived`. The fixture asked to move to
**`triage`**, which U11 deleted. So the route correctly answered `400
Invalid column`, the request never reached `moveTask`, and **every
assertion in the case was unreachable**.

Same class as the two assertions #2720 corrected in
`task-dependency-mutation.pg.test.ts`: a test pinning an id the board no
longer has. It is also **why it sat unclaimed** — the failure message
says *"Invalid column"*, which reads as a broken guard rather than a
stale test, so anyone glancing at it would reasonably assume it belonged
to whoever last touched the move route.

Target changed to `in-progress`: keeps the case's actual subject intact
(a caller-supplied `bypassGuards` / `moveSource` must not be forwarded)
and is a forward move from `todo`, so the R16 backward-move PR guard
stays out of the way. **The point was never which column.**

## The paired cases the suite lacked

The cheapest wrong fix here would have been to relax the route's
validation until the old fixture passed again. Two cases now make that
impossible:

- an **undeclared** column (`triage`) is still **rejected**, with a
message naming the board's own columns so an operator can act on it;
- a **declared** column is still **accepted**, so validation cannot
degrade into "reject everything".

Without them, the next person to see *"Invalid column"* in a failure
cannot distinguish a stale fixture from a broken guard — which is
exactly the half-hour this cost me.

## Verification

11/11 route suites green (**69 tests**, was 1 failing) · `pnpm
test:gate` **487 / 71** · `tsc -p packages/dashboard` clean · `pnpm
lint` clean.

Scope is one test file. Opening this despite the one-PR-per-worker rule
because it is the same category as #2753 — a red on main, which the
close-out brief made priority one — and it touches nothing my other six
PRs touch.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 06:51:10 -07:00
gsxdsm
e0010f241e fleet: github-tracking-comments.ts 9 → 3 — and the sync-filter class is now in a FOURTH file (~25 sites on one decision) (#2715)
Claiming the **github-tracking pair**. This converts the comments half;
the reconciler half is the flagged class, with the evidence below.

## Census before/after

| | before | after |
|---|---:|---:|
| `github-tracking-comments.ts` | **9** | **3** |

Baseline re-recorded; `--strict` exits 0.

## Converted: 6

The `event.to === "in-progress"` / `=== "done"` sites in
`handleTaskMoved`, to the **wip** and **complete** roles. One
resolution, placed **immediately after the tracking-enabled gate** — so
a move on an **untracked** task pays nothing, which is most moves in
most projects.

## Deliberately not converted: 2 — the ordering is the reason

```ts
if (event.to !== "in-progress" && event.to !== "done") return;   // line 232
```

This runs **before** the tracked-task gate. Converting it moves the
resolution ahead of that gate and makes **every task move in the
project** resolve a workflow just to decide the task has no GitHub
issue. That's a real cost on the hottest event in the system, to convert
a guard whose only job is a cheap filter. Recorded at the site.

## The remaining 1

**Line 165** — `transition === "done"` inside `formatTrackingComment`, a
**pure formatter** with no store and no task. Same shape as
`project-engine.ts:2555`. Threading a resolution into a formatter to
pick a string is the wrong trade.

## `github-tracking-reconciler.ts` (9) — not claimed here, and here is
why

All nine are:

```ts
.filter((task) => task.column === "done" || task.column === "archived")
```

Synchronous filters over task **lists**, where per-task resolution is N
awaits inside a sync predicate.

**This is the fourth file with that exact shape** — after `store.ts`
(#2709), and the dependency pairs in `TaskDetailModal` (#2696) and
`register-task-workflow-routes` (#2700). By my count **roughly 25 sites
across four files now wait on one decision**:

1. **Prefetch lifecycle columns alongside the task list** and pass a
resolved map into these predicates — keeps them synchronous, one
resolution per *distinct workflow* rather than per task. This is the one
I'd argue for.
2. Make the predicates async and accept per-task resolution.

It's a design change, and four workers guessing separately is exactly
how two halves of one rule drift apart. One decision covers all of them.

## Verification

`pnpm test:gate` **GREEN** (158 + 10 + 487 + 71) ·
`github-tracking-comments` + `github-issue-comment` **81/81** · `pnpm
lint` clean · dashboard `tsc` clean · `--strict` exits 0.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 06:39:12 -07:00
gsxdsm
72d42652e5 fleet: CLI surface 16 → 0 — 'active=0' on a busy board, and a retry gate that disagreed with the dashboard (#2728)
**Claimed on #2714 before starting.**
`packages/cli/src/commands/task.ts` (8) + `dashboard.ts` (8) — **16 →
0**.

## The finding that matters: `active=0` on a busy board

The same four-line aggregation appears **four times** in `dashboard.ts`
— the TUI stats refresh, the serve summary, the status line, the
agent-stats pass. Each compared the default lineage's two ids, so on a
renamed board every one reported `active=0` while the board was plainly
busy.

**This is worse than an inert internal guard.** A recovery path that
silently stops firing is invisible until something breaks. A stats line
that says zero is **read, believed, and acted on** — *"nothing is
running, so I can restart the engine."*

The four copies are now one helper, and that is the other half of the
fix: four independent copies of a lifecycle decision is how they drift,
and these were identical **by accident, not by construction**. One IR
read per *workflow*, asserted by call count — because the returned
number is identical either way, so only counting the work can see it.

## The retry gate exists twice, and #2713 converted one of them

After #2713, `POST /tasks/:id/retry` accepted a renamed board's stalled
review card while `fn task retry` refused it with *"not in a retryable
state"* — **one operator action answering differently depending on the
surface**.

The rule, stated at the site: **converting one copy of a duplicated gate
creates a disagreement that is harder to diagnose than the original
inert guard.** Grep the classifier by name before calling a lane
converted.

## The rest

- **`fn task set-node` / `clear-node`** rewrote the node override of an
*actively executing* card, because the "is in progress" check never
matched. That guard exists because the rewrite races the run.
- **The duplicate-guard candidate filter** kept completed cards in the
comparison set on a renamed board, so a new task was reported as a
duplicate of work that had already landed — the opposite of useful.
- **The duplicate-lineage `(archived)` marker** never printed, so the
operator could not tell a live duplicate from a filed one.

## Two DELIBERATE-LITERALs, with reasons

The board-render glyph compares `col` taken from the legacy `COLUMNS`
enum **that loop iterates** — the literal matches its own receiver by
construction. The real defect is already named in the code above it: a
card in a renamed column **is not rendered at all**, which is the R8/U10
surface change, not this glyph. Converting it would hide that behind a
trait lookup while the loop still cannot see the card.

## Pre-existing, not mine

5 failures in `commands/__tests__/task.test.ts` (GitHub import) **fail
on `origin/main`** — verified by stashing this change and re-running.
Someone owns that; it should not ride in here.

## Verification

census **16 → 0** · `pnpm test:gate` **10 / 71** · `pnpm smoke:boot`
**PASS** · `tsc -p packages/cli` clean · `pnpm lint` clean ·
`task-retry` 3/3 · 4 new cases with **2 red on revert**.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 06:04:23 -07:00
gsxdsm
5791dfeeb7 fix(dashboard): the routes and the engine disagreed about what "review" is — a converted guard that still contradicts core (#2723)
Built **on top of #2713**, which converted this file to trait-resolved
membership sets while my #2701 was open on the same file. Their design
is better than the single-id resolver I had — the arity argument in
their comment (a single id answers "where should this card GO", a SET
answers "is this card ALREADY there") is correct, and I have closed
#2701 rather than contest it.

Two things survive that #2713 did not cover.

## 1. The routes and the engine disagreed about what "review" IS

Core's `resolveLifecycleColumns().review` — the answer the **executor**,
the **merger** and **project-engine** all act on — resolves review from
`mergeOrchestration`. This file's resolver looked only at `mergeBlocker`
/ `humanReview`.

The default lineage hides it, because `in-review` carries all three:

```
in-review[merge-blocker, human-review, stall-detection, merge]
```

A board that declares only `merge` on its review lane resolved as
**review in the engine** and **not review in the routes**. The executor
treats the card as in review; comment re-engagement, the retry gate and
branch-binding recovery all refuse it.

**That is the same defect class as a literal, one level up:** the guard
is converted, it reads a real trait, and it still disagrees with the
authority. Unioning all three makes this resolver a **superset** of
core's, so the routes cannot refuse a card the engine considers in
review.

How I found it is worth stating: my renamed-board fixture declares only
`merge` on its review lane, and the ported suite failed against main. I
nearly "fixed" the fixture to add `mergeBlocker` — which would have
papered over a real cross-layer inconsistency to make my own test pass.

## 2. The 400s name the board's own columns

#2713 converted both gates but left the messages saying `in-review` /
`in-progress`. **Being told your card must be in a column that does not
exist is worse than a wrong guard** — a wrong guard is a bug report; a
wrong column name sends the operator looking for something that was
deleted.

## Tests

7 cases ported from #2701 and adapted to the membership-set shape. **3
red on revert** of the `mergeOrchestration` union.

## A pre-existing failure I am reporting, not folding in

`register-task-workflow-routes.move-bypassguards.test.ts` **fails on
`origin/main`** — 400 where it expects 200. Verified by stashing my
change and re-running, so it is not mine. Someone owns that regression;
it should not ride in on a lane-consistency PR.

## Verification

`pnpm test:gate` **10 / 71** · the ten other
`register-task-workflow-routes` suites pass · 7/7 new · `tsc -p
packages/dashboard` clean · `pnpm lint` clean · census `--strict` exit
0, no count change (this PR converts nothing new — it makes an
already-converted resolver agree with core).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 05:25:28 -07:00
gsxdsm
a59c6aea50 fleet: register-task-workflow-routes.ts 20 -> 5 (#2713)
## Census before / after

| | before | after |
|---|---:|---:|
| `register-task-workflow-routes.ts` | **20** | **5** |

I handed this cluster off in an earlier turn with an analysis rather
than a conversion. Coming back to it with the analysis already done made
it tractable.

## Converted with the file's own idiom

This file already had `resolveIntakeColumnForTask` /
`resolveWipColumnForTask` / `resolveReboundColumnForTask`: resolve the
column **id** from the task's workflow, fall back to the legacy id when
the IR cannot be read. I added the three roles it was missing — review,
complete, archived — in that same shape.

**Deliberately not the `columnRoles` predicate helpers** used in
`packages/dashboard/app`. Those take resolved trait *flags*, which a
route handler does not have — it has a store and a task id, and must do
an async lookup. Importing them here would mean fetching flags per
request to answer a question this file already answers a simpler way.
One idiom per layer.

Every converted handler **resolves once and reuses**, so two checks in
the same request cannot disagree about which column is the review lane.
The retry handler had three separate review checks and now shares one
resolution.

## Also fixes an inversion

`isArchived` ORed the legacy id with the resolved trait
**unconditionally**, so a column merely *named* `archived` counted as
archived even when its own workflow said otherwise. Same pattern
previously found in `TaskContextMenu` and `isPreExecutionHoldColumn`.
Now flags-first, id as fallback.

## The five survivors, each with a reason

| count | line | why |
|---:|---|---|
| 1 | 1193 | `tasks.filter(t => t.column === "todo")` — a **list** path.
Per-task IR resolution is N store reads on board load; the site already
carries a note measuring that cost. Needs the hold column resolved per
*workflow* from the board payload — a real change, not a rename. |
| 1 | 1926 | **Already flags-first**: `moveTargetIr && declaresColumns ?
columnHasFlag(...) : column === "in-progress"`. The literal *is* the
documented no-IR fallback. |
| 1 | 4872 | The fallback arm of the `isArchived` fix above — same
shape, deliberately kept. |
| 2 | 4024 | `depTask.column` — a **dependency's** column, not the
task's. Resolving it means fetching that task's IR per dependency; out
of scope. |

So three of the five are correct as they stand, and two are genuinely
deferred.

## Verification

`plan-approval-intake-column` + `stranded-refinements-routes` 13/13.
`pnpm test:gate` green (10 / 158 / 487 / 71). `pnpm
check:lifecycle-columns` exits 0 with the baseline re-recorded here.
`tsc -p packages/dashboard/tsconfig.json` clean. `pnpm lint` clean.

Typecheck was run after **every** batch rather than once at the end —
with 15 edits across async handlers in a 6000-line file, a single late
typecheck would not tell me which batch broke it.

No changeset: no user-visible change.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 04:17:03 -07:00