## Summary
- Advance the matched Pi runtime pin (`pi-ai`, `pi-coding-agent`,
`pi-agent-core`, `pi-tui`) from **0.82.0 → 0.82.1** so
electron-builder's production-dependency walk accepts `pi-agent-core`'s
`pi-ai@^0.82.1` requirement.
- Fixes the Desktop packaging PR-lane failure:
`Production dependency @earendil-works/pi-ai not found for package
@earendil-works/pi-agent-core` (required `^0.82.1`).
- Keep the workspace override guard; update pin-policy fixtures and CLI
package-config expectations.
- Tighten the advisory packaging step-order test so it asserts against
the real `electron-builder --dir` step (not a missing release-only step
name that previously passed via `indexOf === -1`).
- Run `pnpm dedupe` so the packaging lane's lockfile dedupe
early-warning is clean.
## Context
#2439 pinned the full Pi closure at 0.82.0 and made recent main-based
packaging runs green. This advances to the current upstream patch so
deploy + electron-builder stay aligned with `pi-agent-core@0.82.1`'s
declared dependency range.
## Test plan
- [x] `node scripts/check-pi-versions-pinned.mjs`
- [x] `node --test scripts/__tests__/check-pi-versions-pinned.test.mjs`
- [x] `pnpm --filter @runfusion/fusion exec vitest run
src/__tests__/package-config.test.ts`
- [x] `pnpm --filter @fusion/desktop exec vitest run
src/__tests__/release-workflow.test.ts`
- [x] `pnpm dedupe --check`
- [ ] GitHub: Desktop packaging (should run full packaging walk —
lockfile/package.json touched)
- [ ] GitHub: PR Checks (Lint, Typecheck, Build, Gate)
## What
Plan Review, planning, and the replan loop move from the implementation
column into the **planning lane** (`todo`), so a task under
specification never holds a WIP slot. The card crosses into
`in-progress` exactly once, at `parse`, released by the scheduler.
Operators also finally see a **Plan Review** badge while the gate runs —
it was previously invisible on the default workflow.
## The part that made it possible
Moving the node is ten lines. It was attempted three times and reverted
each time, because a graph run with no durable continuation replayed
from `start` and dragged an in-progress card *backward* out of the WIP
column, firing `abort-on-exit` and stranding it in a pre-WIP column with
no releaser.
So this PR adds the graph **entry contract** —
`resolveColumnResumeNode`:
| Card is in | Resumes at |
|---|---|
| `triage` | `start` |
| `todo` | `plan` |
| `in-progress` | `parse` — never re-plans, never moves backward |
| `in-review` | first review node — gates are not skipped |
`ir.columns` is ordered and that order is the lifecycle order; rework
and failure edges are excluded so the entry point is always the main
path. The proof it's the right fix: **`executor-task-done-invariant`
passes unmodified** after failing every previous attempt.
## Also in here
- **Release gate narrowed twice.** `isUnplannedForExecution` applies its
pre-release plan-review gate only when the node's column equals the
card's column *and* the group is enabled for the task. The enablement
check fixes a real deadlock — a task with Plan Review toggled off was
held forever waiting for evidence nothing would ever write.
- **Badge cleanup.** Gate badge reads "Plan Review" instead of the
ambiguous "Reviewing" and no longer hides behind a lane restriction; the
status badge stops duplicating it; `planning` renders as "Planning"
instead of the raw engine token.
- **Coding (Ideas)** renames its planner column to "Planning" (id `todo`
unchanged) and loses its private planning-node re-home — the graph it
clones is already plan-in-place.
- **New sweep** `reconcileUndeclaredTaskColumns` re-homes a row whose
column its workflow no longer declares. Written for a follow-up, kept
because it makes any column edit survivable.
## Test changes
Scheduler and release fixtures now model a card whose Plan Review passed
— the state every real card is in when the capacity sweep sees it. A
held unreviewed card is the gate working, and that path stays owned by
`pre-release-plan-review.test.ts`.
New `workflow-graph-entry-contract.test.ts` covers the invariant at
every lifecycle position, plus the gap-column and remediation-node
cases.
## Verification
Gate 299 + 70 + 10, dashboard badge suites 672, engine
workflow/entry/executor suites 147, core 122. Lint and typecheck clean.
Full engine suite sits at the pre-existing baseline (notifier /
plugin-runner / notification-service, untouched by this).
## Follow-up
Removing the Todo column entirely is a separate ~207-site
lifecycle-vocabulary refactor — planned in
`docs/plans/2026-07-26-001-refactor-workflow-owned-lifecycle-plan.md`
(companion docs PR).
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Plan Review now runs in the Planning lane before implementation
begins.
* Cards resume from their current workflow column without replaying
earlier steps.
* Added automatic recovery for cards stranded in outdated workflow
columns.
* **Improvements**
* Renamed the Coding (Ideas) planner column to “Planning.”
* Refined Plan Review gating to respect enabled settings and the card’s
current column.
* Updated planning and Plan Review badges for clearer, consistent labels
across cards and lists.
* **Bug Fixes**
* Improved workflow transitions and release behavior around planning,
review, and execution.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
## Summary
- reconcile successful graph-native workflow results with pending task
checklist steps even when review handoff already moved the card into the
merge column
- preserve terminal, paused, and no-redundant-move behavior
- cover the real Compound Engineering post-review-handoff state with a
regression test
## Root cause
Compound Engineering runs `review-handoff` before `merge`. Review
handoff moves the task to `in-review`, which is also the merge column.
`ensureWorkflowMergeBoundaryTask()` returned immediately for cards
already in that column, before projecting successful
`workflowStepResults` onto legacy `Task.steps[]`. The merger then saw
`0/N` and rejected approved work with `task has incomplete steps`.
## Verification
- RED: regression test failed before the fix because `store.updateTask`
was never called
- GREEN: `executor-graph-boundary.test.ts` — 6 passed
- relevant non-PostgreSQL set — 31 passed, 5 PostgreSQL tests explicitly
skipped
- `@fusion/engine` typecheck passed
- changeset format passed
- `git diff --check` passed
## Baseline note
`ce-workflow-step-executor.test.ts` currently has three failures on
clean `origin/main` after FN-8601 foreach-proof hardening. The same
failures reproduce without this patch and are not regressions from this
change.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Improved reconciliation after review handoff by projecting completed
step results onto the legacy checklist when reaching the merge column.
* Prevented tasks from being marked approved with incomplete step counts
(including “0/N” style states).
* Reduced unnecessary merge failures and deadlock/pause scenarios when
merge-column progress was already recorded.
* **Tests**
* Added coverage for execute-and-merge workflows, ensuring
merge-boundary resolution updates pending steps without moving the task.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Deletes two pieces of automated "meta" machinery that filed and
garbage-collected cards restating state already on the task that failed.
Net **-1015 lines**.
## Why
**Automated recovery follow-ups.** `createAutomatedFollowup` and its
dedup engine (289 lines of signature matching, 1h recurrence
rate-limiting, 24h supersedes windows) existed to file recovery cards
for verification-cap and merge-conflict give-ups. In both cases the
parent is *already* parked `failed` with a descriptive `error` and a log
entry carrying the failing command, branch, and output — the card was a
second copy of that.
**Meta-task auto-archive.** The sweeps that garbage-collected those
cards were worse than redundant: the regex classifier matched ordinary
feature work, and its positional fallback bound cards to unrelated
tasks, so **live work could be archived**.
They are removed together, because the auto-archive sweeps only existed
to clean up after the follow-up engine.
## What changed
### Deleted
- `packages/engine/src/verification-followup-dedup.ts` in full —
`createAutomatedFollowup`, `decideAutomatedFollowup`,
`AutomatedFollowupKind`, `computeVerificationFailureSignature`,
`extractFailingTestFiles`.
- `findActiveRecoveryFollowUp` — dead code, defined and never called
(`tsc` independently flagged it `6133 declared but its value is never
read`).
- The meta-task auto-archive sweeps `autoArchiveResolvedMetaTasks` /
`autoArchiveStalledMetaTasks` and helpers `classifyMetaTask` /
`resolveMetaTargetTaskId` / `computeMetaChainDepth` / `archiveMetaTask`
/ `evaluateMetaAutoArchiveGuards`, plus settings
`metaTaskStallAutoCloseMs` and `metaTaskActiveExecutionGraceMs`.
- Run-audit types `task:auto-archived-meta-resolved`,
`task:auto-archived-meta-stalled`,
`task:auto-archive-meta-resolved-skipped`,
`task:auto-archive-meta-stalled-skipped`,
`verification:followup-created`, `verification:followup-deduped`.
The two signature helpers were **deleted rather than relocated** — once
the three call sites went they were provably unreachable:
`buildVerificationFailureSignature` had exactly one caller, and it was
the only caller of `extractFailingTestFiles`.
### Call sites 1 and 2 — park kept, card dropped
Verification-cap and merge-conflict give-ups keep their park, audit
event, operator comment, and log entry. Site 1's `error` string was
reworded off `"See follow-up task for investigation."` (no follow-up
will exist) to carry the guidance itself. `autoResolveDisabled` was
**kept** — it still drives the outer park guard and the `reason` string;
only the inner branch that guarded card creation is gone.
### Call site 3 — autostash orphan, replaced not deleted
This one is a genuine data-loss guard, so it keeps a durable trail. A
`live`-classified orphan is a merger stash holding **real uncommitted
work**, and unlike sites 1–2 there is no parked parent — the parent may
already be `done` and merged, so nothing else on the board would ever
mention the stash.
The card is replaced by a `logEntry` **and** an `addTaskComment` on the
parent, preserving every fact the old description carried: the sha,
`record.label` (the handle `git stash` recovery needs),
`record.detectedByTaskId`, and `sourcePhase`. New truthful run-audit
event `task:autostash-orphan-live-detected` replaces the borrowed
`verification:followup-*` name, with ids/outcomes-only metadata per
AGENTS.md.
### Kept unchanged: the two real product features
Eval follow-ups (`eval-followups.ts`) and PR-comment follow-ups
(`pr-comment-handler.ts`) only borrowed the shared engine for its dedup
pass. Both keep their exact behavior, column, priority, `sourceType`,
and log lines, with dedup inlined as a `listTasks` scan on
`suggestionId` / `prNumber` respectively. Both fail open (create) if the
listing throws, matching the old engine.
## Test changes — read this one
Two tests asserted the *deleted* engine's rate-limited `"[verification
recurrence]"` logEntry. Those assertions were removed, **not loosened**:
both tests still assert no duplicate card is created, and the eval test
still asserts the existing id is reported back. No coverage of surviving
behavior was weakened. The three `meta-*` test files were deleted along
with the sweeps they covered.
## Verification
```
$ pnpm test:gate
Test Files 2 passed (2) Tests 10 passed (10) # core
Test Files 16 passed (16) Tests 299 passed (299) # engine-core
Test Files 1 passed (1) Tests 70 passed (70) # ci-shape
GATE_EXIT=0
$ pnpm --filter @fusion/engine --filter @fusion/core exec tsc --noEmit -p tsconfig.json
TSC_EXIT=0 (no output)
```
Plus a file-scoped run over the touched surfaces (`eval-followups`,
`pr-comment-handler`, `merger-autostash-orphan-surface`,
`merger-autostash-cleanup`, `run-audit`, `run-audit-secret-taxonomy`,
`project-engine`, `project-engine-manager`): **213/213 passed**.
A repo-wide grep confirms no surviving references to any deleted symbol,
module, or audit event.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Failed tasks now retain recovery and verification details directly on
the original task instead of generating separate follow-up cards.
* Live autostash issues now preserve stash information in task comments
and activity logs.
* Existing evaluation and pull-request follow-ups continue to be reused
when appropriate.
* **Changes**
* Removed automatic archival of meta-tasks.
* Removed obsolete meta-task timing settings.
* **Documentation**
* Updated architecture and settings documentation to reflect these
workflow changes.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
## Summary
- query Task Detail modal assertions from the document after its
FloatingWindow portal migration
- sync secondary locale catalogs with the current English key structure
## Test plan
- `pnpm --filter @fusion/dashboard exec vitest run
app/components/__tests__/TaskDetailModal.attachments-and-tabs.test.tsx
--silent=passed-only --reporter=dot`
- `pnpm i18n:status`
- `pnpm --filter @fusion/dashboard build`
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Documentation**
* Added missing localization entries for workflow selection/help text,
action-based report targeting, agent tool output limit hints,
auto-update/restart messaging, and task refinement status labels across
Spanish, French, Korean, Simplified Chinese, and Traditional Chinese.
* **Tests**
* Updated task detail modal tests to query the correct rendered document
root so portal/modal content is asserted reliably.
* Minor test formatting adjustments to keep assertions consistent
without changing coverage.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
## Summary
- update the touch-resize browser fixture assertions for New Task's
FloatingWindow migration
- verify the shared nine touch targets, resize handle ID, and unified
geometry persistence key
## Test plan
- `corepack pnpm --filter @fusion/dashboard build`
- `corepack pnpm --filter @fusion/dashboard exec vitest run --project
dashboard-browser-touch --silent=passed-only --reporter=dot
src/__tests__/task-modal-touch-resize-browser.test.ts`
- `corepack pnpm check:changesets`
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Tests**
* Updated tablet touch-resize coverage for task-detail to use the shared
southeast resize handle and corrected expected hit-target counts.
* Refreshed hit-target detection and drag target assertions after
switching the test surface to the task-detail/floating window layout.
* Updated persistence/geometry validations to compare the stored unified
width/height against the resized panel dimensions.
* Reworked desktop vs tablet assertions, including tighter overlay
padding checks and standardized shared resize-handle sizing
expectations.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
## Summary
- add the missing `debug` method to the runtime-resolution logger mock
- prevent logger calls from short-circuiting runtime selection
assertions
## Test plan
- `pnpm --filter @fusion/engine exec vitest run
src/__tests__/runtime-resolution.test.ts`
- `pnpm --filter @fusion/engine typecheck`
- `pnpm check:changesets`
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Tests**
* Updated the runtime-resolution test suite’s mocked logger to also
support debug-level messages, alongside existing log, warn, and error
handling.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
## Summary
- restore the two auto-update settings keys in all five secondary locale
catalogs
- keep generated locale structure aligned with the authored English
catalog
## Test plan
- `pnpm i18n:status`
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Localization**
* Added new settings text keys for “automatic updates and restart” (and
its help text) in Spanish, French, Korean, Simplified Chinese, and
Traditional Chinese.
* The new values are placeholders (currently empty), so the UI may show
missing/blank text until translations are completed.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
## Summary
Gives external integrations (command palettes, plugin launchers,
alternate dashboard shells) a supported way to **discover the host UI**
— instead of hardcoding the dashboard's view ids, labels and settings
search terms and hand-syncing them on every release. This is the
read-only metadata slice of the "constrained by a stable host context
and API client" idea in
`docs/proposals/2026-07-01-dashboard-theme-plugin-system.md`, and the
follow-on to #2415 (theme tokens + overlay layering).
Two additions, both inert unless called:
| Endpoint | Returns |
|---|---|
| `GET /api/views` | Every registered built-in view id, in dashboard
order — `id`, English `label`, plus optional i18n `labelKey`, legacy
`aliases` and `internal` flag. |
| `GET /api/settings/sections` | Selectable Settings sections — `id`,
`label`, `labelKey`, `scope`, `group`, `keywords`, `searchableKeys`,
`advanced`. |
Both are read-only, return static project-independent metadata, take no
project id, and are mounted inside `createApiRoutes` so they sit behind
exactly the same `/api` authentication as every other dashboard route —
no more, no less.
## What actually changed — one source of truth
The endpoints are the small part. The core of the diff is **collapsing
duplicated UI metadata into two shared registries that now drive both
the dashboard UI and the API**:
- `packages/dashboard/src/shared/dashboard-views.ts` — canonical view
ids + English labels + i18n keys + legacy aliases.
- `packages/dashboard/src/shared/settings-sections.ts` — canonical
settings sections + scope/group/search metadata, with `group` and
`advanced` derived from the list's own structure.
`LeftSidebarNav`, `SettingsModal` and `useViewState` were rewritten to
consume those registries instead of carrying their own copies (net
**−230 lines** in `SettingsModal` alone). Edit the registry and the
rendered UI and the API move together.
## Drift protection
Being precise about what each test can and cannot catch, because
"no-drift" claims are easy to overstate:
- `left-sidebar-nav-registry-parity.test.tsx` — the one test that
catches drift the registry does not already determine. It **renders**
the sidebar with a recording `t()` spy and pins each entry's translation
key and English fallback to the registry (the sidebar still hardcodes
its keys). It also asserts the rendered destination count equals the
enrolled id list, so a newly added sidebar view fails until it is
enrolled.
- `ui-metadata-sync.test.ts` — pins the Settings navigation list,
advanced-visibility set, persisted view list, reset-key registry and
both endpoint payloads to the registries. Since those consumers are now
*derived* from the registries, these assertions mainly guard against a
future consumer **re-hardcoding** its own copy. Two of them do stand on
their own: each section's served `group` is pinned to the group header
it actually renders under, and no published `labelKey` may resolve to a
non-leaf i18n node.
- `register-ui-metadata-routes.test.ts` — drives the real Express router
and asserts each endpoint serves the registry payload verbatim, with no
filtering or reshaping.
- Exactly two **existing** tests are updated, both for the same reason:
they asserted that `SettingsModal.tsx`'s *source text* contains a
section literal that now lives in the registry.
`VoiceInputSection.modal-visibility.test.tsx` now asserts Voice Input's
Basic-mode contract against `SETTINGS_SECTION_METADATA`, and
`mcp-documentation.test.ts` reads the registry for the two MCP section
ids. No other existing test in the package changes.
## Design notes / decisions for review
- **`GET /api/views` returns the full registry, not the live menu.** It
includes flag-gated / experimental ids and `internal` (non-navigable)
destinations; reachability depends on flags and plugins this endpoint
does not evaluate. Documented as "known view ids", not "visible nav
entries".
- **`labelKey` is optional and best-effort; `label` is the guarantee.**
A `labelKey` is published only where the dashboard itself renders that
view's title through it. `graph` (labelled from a plugin manifest) and
the internal `task-detail` carry none rather than advertise a key that
resolves to nothing — and `task-detail` in particular must not point at
`taskDetail.title`, which is an occupied i18n *namespace* whose lookup
returns an object rather than falling through to a default. A guard test
now enforces that. Separately, a few published keys (`nav.ideation`,
`nav.importTasks`, `nav.automations`, `pr.view.title`) are the
dashboard's real keys but aren't in the shipped catalogs yet because the
host supplies their English inline; the docs say plainly that consumers
must fall back to `label`.
- **`keywords` / `searchableKeys` are explicitly non-contractual.**
`searchableKeys` exposes the raw i18n translation-key strings backing a
section's searchable copy; values, ordering and presence may change
between releases. Documented as best-effort search hints, never stable
identifiers.
- **Migration is deliberately partial.** The desktop sidebar, Settings
navigation and persisted view list now come from the registries;
`Header.tsx` and the mobile More sheet still hardcode a few of the same
labels. They can still drift from what `GET /api/views` reports;
converting them is left to a follow-up so this diff stays reviewable.
- **No project scoping, deliberately.** The proposal doc rightly pushes
plugin traffic through a project-scoped client — these two endpoints are
the exception that proves the rule: they return static registry metadata
that is identical for every project, so threading a `projectId` would
imply a scoping guarantee that does not exist here. They never touch
`getScopedStore` / `TaskStore`.
- **Two endpoints rather than one `/api/ui-metadata` envelope.** Views
and Settings sections are independent registries with different
consumers, and `/settings/sections` sits naturally beside the existing
`/settings/*` routes. A consumer that only needs navigation doesn't pay
for settings metadata.
- **The registry extraction ships with the endpoints rather than as a
separate PR.** The registries *are* the mechanism that keeps the API
honest — split apart, the first half is a refactor with no observable
effect and the second can't land without it.
- **Placement:** `packages/dashboard/src/shared/` is a new directory,
and these are the first *production* `app/ → src/` imports in the
package (today the only one is in `ProviderIcon.test.tsx`). They sit
under `src/` because `src/`'s tsconfig cannot import `app/`, so a module
both sides consume has nowhere else to go; both registries are
dependency-free data leaves, and `vite build` plus
`check-no-node-only-core-imports-in-dashboard` confirm the client bundle
is unaffected. The considered alternative was `packages/core/src` behind
the `dashboard-browser-safe-core-modules.json` allowlist, where
`mobile-nav-primary-items.ts` keeps a destination→labelKey table — these
stayed out of `core` because they are dashboard-owned UI ids, and
because the two tables describe different surfaces (core mirrors the
mobile nav's `nav.skills`/`nav.settings`; this registry mirrors the
desktop sidebar's `header.skillsView`/`header.settings`).
- Ships a `@runfusion/fusion` **minor** changeset (`category: feature`).
Happy to adjust any of the above — shape, placement, or dropping
`searchableKeys` — if you'd rather it landed differently.
## Verification
- Rebased onto `main@26dcccb7c`. Two conflicts, both resolved by
absorbing upstream's work rather than reverting it:
- `SettingsModal.tsx` — upstream's `voice-input` section (and the
`FNXC:VoiceInput` decision comment explaining it stays out of the
advanced-only set) moved into the registry. The registry's section list
is byte-identical to `main`'s `SETTINGS_SECTIONS` (45/45 entries, all
fields), and the registry-derived `ADVANCED_SETTINGS_SECTION_IDS` is
identical to `main`'s hardcoded set (19/19, same order) — both verified
mechanically, not by eye. Upstream's `RUNTIME_*`
hide-uninstalled-runtimes sets are untouched.
- `routes/README.md` — the `mount-sequence` list regenerated from
`CREATE_API_ROUTES_REGISTRAR_MOUNT_SEQUENCE`, so `registerVoiceRoutes`
and `registerUiMetadataRoutes` are both in place and the contract test
passes.
- `DASHBOARD_VIEWS` covers exactly `main`'s `BuiltInTaskView` union,
aliases included, and `BUILT_IN_TASK_VIEWS` reproduces `main`'s 27-entry
array in order (`devserver` still preceding `dev-server` for the
migration path).
- Every one of the 20 sidebar labels the refactor rewrote was checked to
be byte-identical to `main`'s hardcoded fallback, and every `FNXC:`
decision comment displaced by the move was accounted for — all 75 in
`SettingsModal.tsx` and all 11 in `useViewState.ts` survive, relocated
onto the registry entries they document.
- The full `dashboard-app` + `dashboard-api` suites were run at this
commit (**20,706 passing**) and again on unmodified `main@26dcccb7c`,
and the failing-file sets compared: **every file that fails here also
fails on `main`** — nothing regresses. The overlap is environment-driven
(Postgres-backed `*.pg.test.ts`, tests needing built `dist` artifacts,
and `SettingsModalNodeRouting.test.tsx`'s `No "fetchSystemInfo" export
is defined on the "../../api" mock`), none of it touched by this change.
- `tsc --noEmit` clean for both dashboard projects, `eslint` clean on
every changed file, and `vite build` of the client bundle succeeds (the
two pre-existing `@fusion-plugin-examples/claude-runtime` /
`playwright-core` module-resolution errors reproduce on unmodified
`main`).
- Repo gate scripts pass: `check-changeset-format`,
`check-routes-modular`, `check-no-node-only-core-imports-in-dashboard`,
`check-no-cwd-relative-dashboard-test-reads`, `check-mock-completeness`.
- The three new assertions were mutation-tested rather than assumed
load-bearing: breaking the registry's `group` derivation, dropping an
enrolled sidebar id, and re-pointing `task-detail` at the
`taskDetail.title` namespace each make their test fail.
- Local CodeRabbit review over two passes: 3 minor findings, all
addressed (parity projection missing `group`; route tests asserting
partial instead of exact payloads; the `labelKey` guard not covering the
settings registry).
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added authenticated, read-only APIs for discovering dashboard views
and selectable Settings sections.
* Added dashboard view metadata, including labels, aliases, internal
status, and translation keys.
* Added Settings metadata with grouping, scope, advanced status, and
search-related information.
* Updated navigation and Settings UI labels to use shared metadata.
* **Documentation**
* Documented the new metadata endpoints and integration guidance.
* **Bug Fixes**
* Added safeguards and automated checks to keep UI navigation and API
metadata synchronized.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Co-authored-by: Claude <noreply@anthropic.com>
## Summary
- refresh the engine synchronous-shellout allowlist after recent
self-healing and executor source additions shifted audited call sites
- keep the guard's path, primitive, and signature checks unchanged
## Test plan
- `corepack pnpm --filter @fusion/engine exec vitest run
src/__tests__/engine-no-blocking-shellout.test.ts --silent=passed-only
--reporter=dot`
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Tests**
* Updated static validation allowlists for audited synchronous shell
command operations so matching stays accurate with the latest call-site
locations.
* Kept safeguards that prevent unapproved blocking shell commands from
passing validation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Co-authored-by: gsxdsm <gsxdsm@users.noreply.github.com>
Two task origins had no workflow picker in front of the operator and always
inherited the project default: `fn task create` (CLI + the `fn_task_create`
agent tool) and refinement tasks. Add a Project General setting for each, where
blank/unset means "Selected workflow" (the operator's current Board lane,
falling back to the project default) and a concrete id pins that origin.
Because the Board lane lives in browser localStorage, non-browser callers could
not resolve "Selected workflow" at all. `boardSelectedWorkflowId` mirrors the
lane into project settings so they can. Note this makes the mirrored lane
project-scoped: two operators on one project share it, last switch wins. The
Board never reads it back, so the only effect is which workflow a newly created
task inherits.
Resolution is `TaskStore.resolveOriginWorkflowOverrideId(origin)`: pinned
setting -> mirrored lane -> `undefined` to inherit each caller's existing
default-workflow path unchanged. A deleted or fragment id degrades to inherit
rather than throwing, so a stale settings value can never break task creation.
An explicit `workflow_id` argument to `fn_task_create` still wins.
Separately, a refinement is now titled by the operator's own feedback via the
shared `deriveFallbackTaskTitle`, not `Refinement: <parent title>`. Ten
refinements of one task previously rendered ten identical titles, so the board
could not tell them apart while the text saying what each one asked for sat in
the description. Provenance moves to a `Refines <id>` card chip alongside the
existing detail-view parent link and dependency edge.
Verified: merge gate (299 tests), lint, full build, and typecheck for core, CLI,
and dashboard all pass. New coverage: origin resolution across both origins and
the full precedence ladder, the two settings pickers, the board-lane mirror,
refinement titling (including sibling distinctness), and the card chip.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Concurrent dashboard work added runtime exports that hardcoded module
mocks did not expose, so any suite rendering the affected component threw
"No <export> is defined on the mock". Adds the missing `../api` exports
(system-info probe, update install/restart, cloudflared, provider key and
login helpers, git remotes/branches) and the `isTabletTouchViewport`
viewport helper across the 26 suites that mock those modules.
Verified: all 26 files pass (1012 tests).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
TaskCard inferred "unplanned" from steps.length === 0 while triage's
todo-discovery and the scheduler's dispatch filter both decide from
PROMPT.md seed-ness, so the badges disagreed with the engine in both
directions: a real spec that parsed to zero steps read as "Queued to
plan" while the scheduler already treated it as a WIP-slot candidate, and
a re-seeded card still carrying old steps read as "Ready" while triage
was about to plan it. Either way the badge sent operators to the wrong
cap.
Adds the shared isTaskAwaitingPlanning predicate (replan park, missing
spec, seed-vs-real content) used by both triage's discovery and a new
best-effort `awaitingPlanning` enrichment on GET /api/tasks. TaskCard
derives both badges from that one value — strict complements — and keeps
the step count only as a fallback for SSE payloads and older servers.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sweeping the other views for the FN-8606 typing bug found no further breakage,
but it did expose why the bug shipped: almost all field coverage uses
fireEvent.change, which sets a value in one shot on a node it already holds and
never needs the input to stay mounted. A remount is invisible to it.
Adds expectStableTyping (types character by character via userEvent, then asserts
DOM node identity, accumulated value, and retained focus) and applies it to the
uncovered surfaces: the FN-8606-migrated AddNode, ConnectNode, NodeDetail,
Scripts, and WorkflowAddStep modals, plus the board's QuickEntryBox composer and
SubtaskBreakdownModal title editing. All pass — this is a detection floor, not a
fix.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Adds fusion-react/no-nested-component-definitions, a custom rule in the house
style of the existing detached-spawn guard. A component declared in render is a
new element type every render, so React remounts its subtree on each parent
update and destroys focus, scroll, and local state.
This pattern shipped three times without review or tests catching it: FN-8606's
ModalShell left Planning Mode and Settings untypable, and MailboxModal's
ReplyContextExpandable collapsed expanded reply rows. Tests missed it because
fireEvent.change sets a value without needing the node to stay mounted.
The rule reports PascalCase functions (including memo()/forwardRef()-wrapped)
that return JSX and are declared inside another JSX-returning function.
Lowercase render helpers are deliberately allowed — they are the sanctioned fix.
Escape hatch: // nested-component-allowlist: <reason>.
Scoped to production .tsx, with a vitest guard for the rule itself. Hoists the
two pre-existing violations (ProviderStatusBadge, GitHubStatusBadge in
ModelOnboardingModal) to module scope so the rule lands clean at "error".
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
evictStaleProcessing cleared `processing` but left the task in
`coordinatorAdmittedTaskIds`, which is only cleared by specifyTask's
finally — the path a hung promise never reaches. The card stayed
eligible (so the throttle branch never logged or emitted
`task:plan-admission-throttled`) while admitOldest's refresh filtered it
out, leaving it on the "Queued to plan" badge with free slots and no
diagnostic until engine restart. Also drop an untransferred pre-held host
slot, which otherwise waits out the 600s stale-excess valve.
Regression tests assert the invariant on the real production candidate
source: an evicted card is re-offered and its host slot returned, while a
still-live stale task keeps both claims.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
ReplyContextExpandable was declared inside MailboxModal's render, making it a new
element type on every render. Any parent update remounted the whole recursive
reply thread, so expanding one reply row collapsed the others and discarded their
DOM identity, focus, and scroll position.
Hoist it to module scope and pass the parent's reply state and handlers through an
explicit env prop so the element type stays stable. Adds a regression test that
expands one row, expands a second, and asserts the first keeps both its node
identity and aria-expanded state — same defect class as the FN-8606 ModalShell fix.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
FN-8606 declared the window shell as a component inside PlanningModeModal and
SettingsModal render. A component declared in render is a new element type on
every render, so React remounted the whole subtree on each keystroke, destroying
the focused input: Planning Mode dropped everything after the first character and
Settings text fields did the same.
Replace ModalShell with a plain renderModalShell(children) call so the returned
element types stay stable, and add a Planning regression test that types
per-character across both the modal and embedded surfaces (fireEvent.change
cannot observe this class of bug, which is why it shipped).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`findLatestByDedupeKey` read `targetContext` through the string-only `fromJson`.
In backend (PostgreSQL) mode that column is jsonb and Drizzle returns it ALREADY
PARSED, so the dedupe scan never matched: every gate retry minted a duplicate
approval request, and an approved grant could never be redeemed. The live
database shows the signature plainly — 17 approved requests, 0 completed.
Normalize both shapes in one place (`normalizeTargetContext`), applied at
`rowToRequest` and both dedupe scan sites, so a row resolves whether it arrives
as a JSON string (SQLite) or a parsed object (Postgres).
The regression test asserts shape-independence rather than the single reported
case: the same stored key must resolve in BOTH shapes, and must not match a
different key or an absent context in either. Mutation-checked — reverting the
scan sites fails exactly the parsed-object case.
Cherry-picked ahead of #2457, which carries the wider approval/permission
hardening pass, because this one is an active production defect on its own.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Keep successful verification responses quiet and return bounded, high-signal diagnostics for failures without hiding zero-work or green-while-red warnings.
The client bundle aliases `@fusion/core` to the leaf `core/src/types.ts` to
keep Node-only dependencies out of the browser, so a package-root import of
`FUSION_CLIENT_HEADER`/`FUSION_DASHBOARD_UI_CLIENT` typechecked but failed
`vite build`:
"FUSION_CLIENT_HEADER" is not exported by "../core/src/types.ts"
Follow the documented pattern instead of widening the root alias: declare a
`./task-delete-attribution` subpath export, add the matching Vite alias ahead
of the broader `@fusion/core` key (Vite matches in order), register the module
in the browser-safe-core allowlist, and import the subpath from the client.
`task-delete-attribution.ts` has no imports at all, so it is a safe leaf.
`app/utils/detectContentLanguage.ts` already warned about exactly this trap;
the miss was mine for verifying with typecheck, lint and test:gate but not
`pnpm build`, which is one of the four checks CI blocks on.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Three related fixes, all originating from a `[api:error] Request failed`
log line showing a 500 on `GET /api/tasks/FN-8610/runtime-fallback`.
1. Missing/deleted tasks now return 404 instead of 500.
`getTaskImpl` signalled a miss with a bare `Error`, and route catches
only mapped errno `ENOENT` to 404 — a leftover from the file-backed
storage era. In Postgres mode nothing sets an errno code, so every
unknown/missing/soft-deleted/wrong-project read returned 500. Adds a
typed `TaskNotFoundError` (message byte-identical) plus a shared
`task-lookup-error` mapper applied across the task, session-diff,
git/GitHub, workflow and file-workspace route registrars. The same
bare throw existed on both archive-lifecycle delete paths, so
`DELETE /tasks/:id` was affected too.
2. 5xx logs now carry the origin stack.
`rethrowAsApiError` constructed a fresh `ApiError` from the message
and discarded the original, so the `FNXC:ApiErrorDiagnostics`
contract logged the rethrow site rather than the throw site — the
reported log entry had no stack at all. Threads `cause` through the
error factories and walks the chain (bounded, cycle-guarded).
3. Task deletions are attributable, and non-operator deletes notify.
`task:deleted` audit rows recorded `agentId: "system"` for every HTTP
delete, making an operator click indistinguishable from a script or
an agent; the calling agent's task id was accepted by the store and
then never persisted. Adds a `callerKind` union recorded in audit
metadata, tags every delete call site, and stamps a self-reported
`x-fusion-client` header from the dashboard client. When the caller
is `agent-tool` or `api-unattributed`, a best-effort notice is sent
to the operator mailbox; operator and engine deletes stay silent.
`x-fusion-client` is attribution, not authentication — anything can send
it. No delete-blocking, gating or permission logic is added here.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Replace empty backendMode stubs with real AsyncDataLayer paths: archive ID
reservation and isTaskArchivedAsync, orphaned task.json re-import, health
snapshots via checkPostgresHealth, settings/agent memory caches for sync
readers, async builtin prompt overrides, and self-healing audit/health
callers that previously used dead sync SQLite fallbacks.
Pin reconcileOrphanedTaskDirsImpl empty result and getDatabaseHealthImpl
always-healthy sentinel under backendMode so inventory category (e) stays
aligned with production self-healing and health call sites.
Drive real sync-reader stubs that empty-return under backendMode
(merge request, workflow selection/overrides/settings, run audit,
legacy step snapshot, settings/health) so category (e) of the
migration inventory stays pinned to shipped behavior.
Inventory analysis found exactly six read-only legacy openers; pin them in a
structural scan and assert incomplete archive guards stay SQLite-free in
backend mode so new production SQLite construction fails CI.
Completes 3b83282273. The classifier and the self-healing reader landed there but
nothing stamped `leaseNodeId`, so the pre-boot reclaim path was unreachable.
Wiring: InProcessRuntime -> TaskExecutor -> WorkflowGraphTaskRunner ->
WorkflowGraphExecutor, which writes the field onto the pending lease.
The executor takes `getLocalNodeId`, a GETTER rather than a value, because the
runtime resolves the node id asynchronously (a CentralCore read) partway through
start() while `executorOptions` is built earlier in the same method. A snapshot
taken at construction would freeze `undefined` and silently disable attribution
forever -- the failure mode where the feature looks wired, typechecks, and never
fires. Reading it at runner-construction time picks up the resolved id.
With this, a review gate whose session dies to an engine restart is reclaimed on
the next self-healing pass instead of waiting out the 15-minute staleness floor.
Peer-owned and legacy unattributed leases still take the floor, so the
double-dispatch protection multi-node depends on is unchanged.
Adds five classifier cases: own-node pre-boot reclaims; peer-node, unattributed,
own-node-post-boot, and no-identity-supplied all still adopt. Verified the first
is not vacuous -- disabling the branch fails exactly that case (1 failed / 13
passed) and no other.
Verified: tsc clean on core and engine, pnpm lint clean, pnpm test:gate green
(299 + 10 + 70), plan-review-lease + plan-review-single-owner +
self-healing-orphaned-pending-step-results green (27).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Groundwork for FN-8603's remaining ~14-minute wait. Liveness for a pending
review gate is judged purely by a 15-minute staleness floor because a lease
records WHO took it (`leaseOwner` = run id) but not WHERE, and under multi-node
every engine sees every other engine's leases. A fresh-but-unknown lease might
be running on a peer, so the floor was the only safe test -- and a lease left by
this node's own crashed process is indistinguishable from it.
Adds `WorkflowStepResult.leaseNodeId` plus an optional `LocalNodeLeaseIdentity`
argument to `classifyReviewLease`. One narrow new case: a lease stamped with the
caller's OWN node id whose `startedAt` predates the caller's process boot is
provably dead -- the process that could have owned it is gone -- so it
classifies as `reclaim` immediately rather than aging out. Deliberately narrow,
because widening it is a double-dispatch risk: absent (legacy) or peer node ids
keep the floor, and a lease taken by this process after boot is still adopted.
InProcessRuntime.start() resolves the local node id from CentralCore (fail-soft;
on error it stays undefined and floor-only semantics apply) and passes it to
SelfHealingManager. The graph executor stamps the field when deps.localNodeId is
set.
NOT YET WIRED, so this is inert in production and behavior is unchanged end to
end: `localNodeId` is not threaded from WorkflowGraphTaskRunner /
WorkflowTaskRuntime down into the executor deps, so no lease actually carries a
`leaseNodeId` yet. The reader is ready; the writer needs that pass-through
(WorkflowGraphTaskRunnerDeps gains the field, the runner forwards it, and the
runtime supplies this.localNodeId). Stopping here rather than half-threading it.
Verified: tsc clean on core and engine, pnpm lint clean, pnpm test:gate green
(299 + 70), core workflow-step-results suite green.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
FN-8603 sat in-review for ~36 minutes after an engine restart killed its Code
Review session 34 seconds in. It did recover on its own; the cost was latency,
not a terminal park.
Sweep ordering. reconcile-orphaned-pending-step-results PRODUCES the failed
results that recover-failed-pre-merge-steps CONSUMES, but in the periodic
maintenance list it ran ~15 entries after it. A step orphaned in cycle N was
therefore rewritten to failed only after recovery had already scanned, so
nothing re-ran it until cycle N+1. Moved it immediately before its consumer and
removed the now-duplicated later entry. Startup recovery already ordered the two
correctly.
Post-review fix budget. Default raised 3 -> 10 per operator request. Three
passes is below the observed convergence length for the gates this fallback
actually governs -- Browser Verification and custom optional gates -- since Plan
Review and Code Review already resolve to "unbounded" when unset, and exhausting
the budget parks the card for a human. The declaration default and five inline
`settings.maxPostReviewFixes ?? 3` call sites in executor.ts/self-healing.ts had
drifted into separate literals, so raising one alone would have left every
unset-settings path on the old value; they now share the exported
DEFAULT_MAX_POST_REVIEW_FIXES.
Not done, and why. Re-dispatching a restart-orphaned lease immediately at
startup is the change that would close the remaining ~14-minute wait, but it is
unsound as specified: liveness is judged by a 15-minute lease-staleness floor
because leases carry no node attribution, so treating a pre-boot lease as dead
would let one node orphan another node's genuinely running review. Needs a node
id on the lease record first. Left the floor intact.
Verified: tsc clean on core and engine, pnpm lint clean, pnpm test:gate green,
self-healing orphaned-pending-step-results and optional-step-revision suites
green.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>