becaacccc4a1cced96756178ff15191cdb98cce8
13956 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
becaacccc4 |
FN-208: add task history reports to task details
Add a History tab that presents per-stage task reports and persists the underlying report data. - Add task-step report storage, migration, serialization, and execution summaries. - Render per-stage report accordions with localized labels in TaskDetailModal. - Add dashboard, core, engine, and schema regression coverage. Files changed: .changeset/fn-208-task-history-tab.md | 7 + docs/dashboard-guide.md | 1 + .../src/__tests__/postgres/schema-applier.test.ts | 5 +- .../core/src/__tests__/task-step-reports.test.ts | 161 +++++++++++++++++++ packages/core/src/index.gate.ts | 3 +- packages/core/src/index.ts | 3 +- .../migrations/0068_fn_208_task_step_reports.sql | 3 + packages/core/src/postgres/schema-applier.ts | 26 ++- packages/core/src/postgres/schema/project.ts | 1 + packages/core/src/store.ts | 2 +- packages/core/src/task-store/merge-queue-ops.ts | 18 ++- packages/core/src/task-store/persistence.ts | 4 +- packages/core/src/task-store/serialization.ts | 1 + packages/core/src/task-store/task-row-mappers.ts | 2 +- packages/core/src/types.ts | 2 + packages/core/src/types/task/task-core.ts | 3 + packages/core/src/types/task/task-log.ts | 14 ++ packages/core/src/workflows/task-step-reports.ts | 49 ++++++ .../dashboard/app/components/TaskDetailModal.tsx | 37 ++++- .../dashboard/app/components/TaskHistoryTab.css | 161 +++++++++++++++++++ .../dashboard/app/components/TaskHistoryTab.tsx | 119 ++++++++++++++ .../TaskDetailModal.attachments-and-tabs.test.tsx | 18 ++- ...skDetailModal.models-progress-workflow.test.tsx | 46 ++++++ .../components/__tests__/TaskHistoryTab.test.tsx | 84 ++++++++++ .../app/utils/__tests__/taskHistory.test.ts | 98 ++++++++++++ packages/dashboard/app/utils/taskHistory.ts | 174 +++++++++++++++++++++ .../src/__tests__/executor-step-summary.test.ts | 89 +++++++++++ packages/engine/src/agent-tools.ts | 10 +- .../engine/src/executor/create-task-update-tool.ts | 12 +- packages/engine/src/executor/execution-prompt.ts | 2 +- packages/i18n/locales/en/app.json | 46 ++++++ packages/i18n/locales/es/app.json | 48 +++++- packages/i18n/locales/fr/app.json | 48 +++++- packages/i18n/locales/ko/app.json | 48 +++++- packages/i18n/locales/pt-BR/app.json | 48 +++++- packages/i18n/locales/zh-CN/app.json | 48 +++++- packages/i18n/locales/zh-TW/app.json | 48 +++++- packages/i18n/src/resources.d.ts | 46 ++++++ 38 files changed, 1501 insertions(+), 34 deletions(-) Fusion-Task-Id: FN-208 Fusion-Task-Lineage: 62318171-9316-4369-8294-07b3b9dd283d Co-authored-by: Fusion <noreply@runfusion.ai> |
||
|
|
c0e99dff10 |
FN-206: collapse task recovery to retry, reset, and delete
Simplify task recovery into three predictable operator actions across the dashboard and lifecycle APIs. - Replace restart-stage recovery with Retry, Reset, and Delete actions. - Update task cards, lists, detail views, menus, translations, routes, and documentation. - Add focused retry behavior tests and remove obsolete restart-stage coverage. Files changed: .changeset/fn-206-removal.md | 7 ++ docs/dashboard-guide.md | 10 +- docs/task-management.md | 2 +- packages/core/src/tasks/task-column-restart.ts | 6 +- packages/dashboard/app/App.tsx | 9 +- packages/dashboard/app/api/legacy.ts | 1 - .../dashboard/app/api/tasks/tasks-lifecycle.ts | 8 -- packages/dashboard/app/components/AppModals.tsx | 3 +- packages/dashboard/app/components/Board.tsx | 6 +- packages/dashboard/app/components/Column.tsx | 5 +- packages/dashboard/app/components/ListView.tsx | 50 +++------ packages/dashboard/app/components/TaskCard.tsx | 82 ++++---------- .../dashboard/app/components/TaskContextMenu.tsx | 81 +++----------- .../dashboard/app/components/TaskDetailModal.tsx | 120 ++++++++------------- .../dashboard/app/components/WorktreeGroup.tsx | 6 +- .../__tests__/RestartStage.host-inventory.test.tsx | 20 ---- .../app/components/__tests__/TaskCard.test.tsx | 35 +++--- .../components/__tests__/TaskContextMenu.test.tsx | 64 ++++------- .../__tests__/TaskDetailModal.rendering.test.tsx | 17 +-- ...etailModal.responsive-and-dependencies.test.tsx | 11 +- .../components/__tests__/TaskDetailModal.test.tsx | 35 ++++-- .../__tests__/TaskRetry.host-inventory.test.tsx | 23 ++++ .../app/components/dashboard/MainContent.tsx | 7 +- .../dashboard/app/components/dashboard/types.ts | 1 - .../app/components/useRightDockController.tsx | 2 - packages/dashboard/app/hooks/useTasks.ts | 6 +- .../app/task-modal-touch-resize-e2e-fixture.tsx | 4 +- .../app/utils/__tests__/taskRetryCopy.test.ts | 27 +++++ packages/dashboard/app/utils/taskRecovery.ts | 5 + packages/dashboard/app/utils/taskRetryCopy.ts | 60 +++++++++++ .../routes-task-retry-stale-merge-status.test.ts | 4 +- ...oute.test.ts => task-retry-stage-route.test.ts} | 96 +++++++++++++---- .../src/routes/register-task-workflow-routes.ts | 63 +++++++---- .../dashboard/src/routes/task-restart-stage.ts | 37 ++++--- packages/i18n/locales/en/app.json | 27 ++--- packages/i18n/locales/es/app.json | 27 ++--- packages/i18n/locales/fr/app.json | 27 ++--- packages/i18n/locales/ko/app.json | 27 ++--- packages/i18n/locales/pt-BR/app.json | 27 ++--- packages/i18n/locales/zh-CN/app.json | 27 ++--- packages/i18n/locales/zh-TW/app.json | 27 ++--- packages/i18n/src/resources.d.ts | 25 +++-- 42 files changed, 589 insertions(+), 538 deletions(-) Fusion-Task-Id: FN-206 Fusion-Task-Lineage: 7ddc5f34-aaf3-4472-8663-6a5d12f40b6d Co-authored-by: Fusion <noreply@runfusion.ai> |
||
|
|
b802d4b3fb |
FN-205: keep the activity feed live and harden resume diagnostics
Keep task activity updates fresh while making resume diagnostics tolerate missing or invalid task identifiers. - Preserve live activity-feed freshness during task hydration and modal updates. - Centralize resume trigger handling and prevent malformed diagnostics requests from returning 400 errors. - Expand dashboard regression coverage and document the updated behavior. Files changed: .changeset/fn-205-feed-freshness.md | 7 + docs/dashboard-guide.md | 2 + docs/diagnostics.md | 6 +- .../dashboard/app/components/TaskDetailModal.tsx | 124 ++++++-- ...TaskDetailModal.feed-stripped-snapshot.test.tsx | 311 ++++++++++++++++++--- .../__tests__/TaskDetailModal.test-helpers.ts | 10 +- .../components/__tests__/TaskDetailModal.test.tsx | 38 +-- .../__tests__/useTasks-hydration-freshness.test.ts | 63 +++++ packages/dashboard/app/hooks/useTasks.ts | 20 +- .../utils/__tests__/resumeInstrumentation.test.ts | 34 ++- .../dashboard/app/utils/resumeInstrumentation.ts | 18 +- .../__tests__/register-diagnostics-routes.test.ts | 39 ++ .../src/routes/register-diagnostics-routes.ts | 15 +- packages/dashboard/src/shared/resume-triggers.ts | 28 ++ 14 files changed, 577 insertions(+), 138 deletions(-) Fusion-Task-Id: FN-205 Fusion-Task-Lineage: dc2e0fd3-741e-4560-91bc-14a31f688d5c Co-authored-by: Fusion <noreply@runfusion.ai> |
||
|
|
d16c8d650b |
FN-204: add task stage restart controls
Add an operator restart action that reruns a task from its current workflow stage. - Add core stage-restart planning and task lifecycle API support. - Expose restart controls across board, list, detail, and grouped task views. - Add localized labels, documentation, and route/core regression coverage. Files changed: .changeset/fn-204-restart-stage.md | 7 + docs/dashboard-guide.md | 8 +- .../core/src/__tests__/task-column-restart.test.ts | 112 +++++++ packages/core/src/index.gate.ts | 10 + packages/core/src/index.ts | 10 + packages/core/src/tasks/task-column-restart.ts | 176 +++++++++++ packages/dashboard/app/App.tsx | 7 +- packages/dashboard/app/api/legacy.ts | 1 + .../dashboard/app/api/tasks/tasks-lifecycle.ts | 8 + packages/dashboard/app/components/AppModals.tsx | 2 + packages/dashboard/app/components/Board.tsx | 10 +- packages/dashboard/app/components/Column.tsx | 5 +- packages/dashboard/app/components/ListView.tsx | 22 +- packages/dashboard/app/components/TaskCard.tsx | 24 ++ .../dashboard/app/components/TaskContextMenu.tsx | 13 + .../dashboard/app/components/TaskDetailModal.tsx | 25 ++ .../dashboard/app/components/WorktreeGroup.tsx | 8 +- .../__tests__/RestartStage.host-inventory.test.tsx | 20 ++ .../components/__tests__/TaskContextMenu.test.tsx | 11 + .../app/components/dashboard/MainContent.tsx | 11 +- .../dashboard/app/components/dashboard/types.ts | 1 + .../app/components/useRightDockController.tsx | 2 + packages/dashboard/app/hooks/useTasks.ts | 6 +- .../app/task-modal-touch-resize-e2e-fixture.tsx | 4 +- .../src/__tests__/task-restart-stage-route.test.ts | 322 +++++++++++++++++++++ .../src/routes/register-task-workflow-routes.ts | 21 ++ .../dashboard/src/routes/task-restart-stage.ts | 166 +++++++++++ packages/engine/src/index.ts | 1 + packages/i18n/locales/en/app.json | 6 + packages/i18n/locales/es/app.json | 6 + packages/i18n/locales/fr/app.json | 6 + packages/i18n/locales/ko/app.json | 6 + packages/i18n/locales/pt-BR/app.json | 6 + packages/i18n/locales/zh-CN/app.json | 6 + packages/i18n/locales/zh-TW/app.json | 6 + packages/i18n/src/resources.d.ts | 6 + 36 files changed, 1044 insertions(+), 17 deletions(-) Fusion-Task-Id: FN-204 Fusion-Task-Lineage: dd8a4493-0db8-41c3-8fe3-cf8ff30e59d5 Co-authored-by: Fusion <noreply@runfusion.ai> |
||
|
|
97fcc5a7ad |
FN-203: support safe workspace task reset
Enable reset to cancel and clean up workspace tasks without deleting the coordinator directory. - Plan per-repository reset targets using workspace layout and ownership checks - Reserve, cancel, and remove workspace worktrees safely while clearing leases and land intents - Document and test workspace reset behavior and publication cleanup Files changed: .changeset/fn-203-workspace-task-reset.md | 7 + docs/dashboard-guide.md | 6 +- docs/workspaces.md | 6 + .../postgres/task-reset-publication.pg.test.ts | 41 +++ .../core/src/__tests__/task-reset-targets.test.ts | 134 +++++++++ packages/core/src/index.gate.ts | 2 + packages/core/src/index.ts | 2 + packages/core/src/task-store/reset-lifecycle.ts | 16 ++ packages/core/src/tasks/task-reset-targets.ts | 103 +++++++ .../task-reset-workspace-lifecycle.test.ts | 306 +++++++++++++++++++++ .../src/routes/register-task-workflow-routes.ts | 200 +++++++++----- 11 files changed, 753 insertions(+), 70 deletions(-) Fusion-Task-Id: FN-203 Fusion-Task-Lineage: ac2406cb-6199-4c6b-82aa-d4790a677e6b Co-authored-by: Fusion <noreply@runfusion.ai> |
||
|
|
6fd4fdd437 |
FN-9217: Fix project-scoped routine and automation listing
Project-scoped routine and automation reads now resolve the engine-backed project stores. - Preserve legacy empty-list behavior for omitted and global scopes. - Return project records through the same stores used for creation and surface unavailable stores as 503 errors. - Add route regressions for populated, duplicate, empty, and unavailable project-store cases. - Add a patch changeset for the published Fusion package. Files changed: .changeset/fn-9217-project-scoped-routine-list.md | 7 ++ .../src/__tests__/routes-automation.test.ts | 121 +++++++++++++++++++-- .../src/routes/register-plugins-automation.ts | 26 +++-- 3 files changed, 134 insertions(+), 20 deletions(-) Fusion-Task-Id: FN-9217 Fusion-Task-Lineage: bb2896d6-3e30-4b84-b86f-4a7c9394603a Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai> |
||
|
|
8fa9acbd6d |
FN-9216: add the Flexoki dashboard theme
Add the Flexoki palette as a fully persisted dashboard theme across web and desktop. - register Flexoki in shared theme types, selectors, startup validators, and preview swatches - define coordinated dark and light token palettes with regression coverage - document theme availability and add a minor release changeset Files changed: .changeset/fn-9216-flexoki-theme.md | 6 ++ docs/dashboard-guide.md | 5 +- docs/settings-reference.md | 3 +- packages/core/src/types/ui/execution-and-ui.ts | 2 + .../dashboard/app/__tests__/flexoki-theme.test.ts | 102 +++++++++++++++++++++ .../dashboard/app/components/ThemeSelector.css | 15 +++ .../components/__tests__/ThemeDropdown.test.tsx | 2 +- .../components/__tests__/ThemeSelector.test.tsx | 2 +- .../__tests__/CommandCenterControls.test.tsx | 2 +- packages/dashboard/app/components/themeOptions.ts | 2 + packages/dashboard/app/index.html | 3 +- packages/dashboard/app/public/theme-data.css | 85 +++++++++++++++++ packages/desktop/src/renderer/index.html | 2 + 13 files changed, 224 insertions(+), 7 deletions(-) Fusion-Task-Id: FN-9216 Fusion-Task-Lineage: cd89ff1a-6b08-4e3d-9103-993f04d583f3 Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai> |
||
|
|
0fd247b589 |
FN-9214: deliver secret values in model-visible tool content
Make policy-approved fn_secret_get reads visible to calling agents without weakening refusal paths. - Include revealed plaintext in tool content for automatic and redeemed approval reads while retaining details.value. - Keep denied, pending, missing, and ambiguous outcomes plaintext-free and audit metadata sanitized. - Add PostgreSQL-backed regression coverage, shared secret-store test setup, documentation, and a patch changeset. Files changed: .changeset/fn-9214-secret-get-value-delivery.md | 7 ++ docs/architecture.md | 2 +- docs/secrets.md | 11 ++- .../__tests__/extension-permission-gates.test.ts | 43 ++------- .../extension-secret-get-value-delivery.test.ts | 101 +++++++++++++++++++++ packages/cli/src/__tests__/pg-extension-harness.ts | 18 +++- packages/cli/src/extension.ts | 17 +++- 7 files changed, 156 insertions(+), 43 deletions(-) Fusion-Task-Id: FN-9214 Fusion-Task-Lineage: 2738aa00-9e67-4282-b950-01ed33431cfc Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai> |
||
|
|
25f24c7439 |
FN-9215: scope skill discovery to selected projects
Resolve skill inventory and catalog installation state against the dashboard's selected project. - create project-rooted package managers for daemon, dashboard, and serve entry points - discover and deduplicate project-local Fusion, Pi, and agent skills with stable precedence - annotate catalog entries as installed and refresh catalog state after installation - document project-scoped behavior and add CLI, route, adapter, desktop, and mobile coverage Files changed: .changeset/fn-9215-project-scoped-skills.md | 7 + docs/dashboard-guide.md | 6 +- .../__tests__/skills-package-manager.test.ts | 50 +++++++ packages/cli/src/commands/daemon.ts | 2 + packages/cli/src/commands/dashboard.ts | 2 + packages/cli/src/commands/serve.ts | 2 + .../cli/src/commands/skills-package-manager.ts | 15 +++ packages/dashboard/app/components/SkillsView.tsx | 10 +- .../app/components/__tests__/SkillsView.test.tsx | 7 +- .../__tests__/skills-view-mobile.test.tsx | 23 ++++ .../__tests__/skills-adapter-project-scope.test.ts | 94 +++++++++++++ .../__tests__/register-agent-skills-routes.test.ts | 146 +++++++++++++++++++-- .../src/routes/register-agent-skills-routes.ts | 8 +- packages/dashboard/src/skills-adapter.ts | 109 ++++++++++++--- 14 files changed, 440 insertions(+), 41 deletions(-) Fusion-Task-Id: FN-9215 Fusion-Task-Lineage: 79da1b10-f317-4df9-af49-1d65b0825c55 Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai> |
||
|
|
5e14551e60 |
fix(docker): ship google-chrome-stable in the runner image
The image shipped no browser at all, silently breaking two features that launch an existing browser and download none: the agent-browser plugin (playwright-core does not fetch a browser at install time) and the Chrome DevTools MCP server. Installs Chrome from Google's signed apt repository, which publishes both amd64 and arm64, so the existing arch pattern resolves on either host. |
||
|
|
906692006a |
FN-202: add workspace merge parity coverage
Verify workspace main-checkout completion handling and multi-repository merge parity through documentation and regression coverage. - Document workspace completion and merge parity expectations. - Cover main-checkout guard wedges, lifecycle parity, and workspace merger behavior. Files changed: docs/settings-reference.md | 1 + docs/workspaces.md | 2 + .../executor-workspace-main-checkout-guard.test.ts | 147 +++++++++++++- .../__tests__/workspace-lifecycle-parity.test.ts | 102 +++++++++- .../engine/src/__tests__/workspace-merger.test.ts | 219 ++++++++++++++++++++- 5 files changed, 464 insertions(+), 7 deletions(-) Fusion-Task-Id: FN-202 Fusion-Task-Lineage: 0e0eda72-40c8-45f8-a767-18c27a6ad6d1 Co-authored-by: Fusion <noreply@runfusion.ai> |
||
|
|
5c111be5f2 |
FN-198: remove task move affordances from dashboard
Remove manual task-movement controls throughout the dashboard while preserving task detail and workflow behavior. - Remove Move to controls and supporting handlers from board, list, card, context-menu, and task-detail surfaces. - Update localized strings, documentation, responsive styling, and affected tests. - Add regression coverage and a release changeset for the dashboard affordance removal. Files changed: .changeset/fn-198-removal.md | 7 + docs/dashboard-guide.md | 6 +- docs/task-management.md | 8 +- packages/dashboard/app/App.tsx | 3 +- .../task-detail-modal-tablet-width.test.ts | 10 +- packages/dashboard/app/components/AppModals.tsx | 1 - packages/dashboard/app/components/Board.tsx | 2 +- packages/dashboard/app/components/Column.tsx | 94 +--- packages/dashboard/app/components/Lane.tsx | 2 +- packages/dashboard/app/components/ListView.tsx | 103 +--- packages/dashboard/app/components/TaskCard.tsx | 125 +---- .../dashboard/app/components/TaskContextMenu.tsx | 222 +-------- .../dashboard/app/components/TaskDetailModal.css | 118 +---- .../dashboard/app/components/TaskDetailModal.tsx | 269 +---------- .../app/components/__tests__/Column.test.tsx | 57 +-- .../app/components/__tests__/ListView.test.tsx | 32 -- .../app/components/__tests__/TaskCard.test.tsx | 286 +---------- .../components/__tests__/TaskContextMenu.test.tsx | 181 ++----- .../TaskDetail.mobile-transition.test.tsx | 4 - .../TaskDetailModal.allow-resurrection.test.tsx | 4 - .../TaskDetailModal.attachments-and-tabs.test.tsx | 44 +- .../TaskDetailModal.create-pr-e2e.test.tsx | 2 - .../TaskDetailModal.create-pr-integration.test.tsx | 2 - .../__tests__/TaskDetailModal.create-pr.test.tsx | 8 - .../TaskDetailModal.custom-fields.test.tsx | 3 - .../TaskDetailModal.definition-actions.test.tsx | 69 +-- ...TaskDetailModal.feed-stripped-snapshot.test.tsx | 1 - ...TaskDetailModal.github-tracking-header.test.tsx | 1 - ...ailModal.github-tracking-renamed-lanes.test.tsx | 1 - .../TaskDetailModal.github-tracking-stale.test.tsx | 7 - .../TaskDetailModal.gitlab-tracking.test.tsx | 2 - ...lModal.inline-editing-and-integrations.test.tsx | 99 +--- ...skDetailModal.models-progress-workflow.test.tsx | 36 -- .../TaskDetailModal.oversight-controls.test.tsx | 40 +- .../TaskDetailModal.oversight-mobile.test.tsx | 18 - .../TaskDetailModal.plan-summary.test.tsx | 3 - .../TaskDetailModal.popup-hidden-gating.test.tsx | 1 - .../__tests__/TaskDetailModal.pr-tab.test.tsx | 3 - .../__tests__/TaskDetailModal.refine.test.tsx | 1 - .../__tests__/TaskDetailModal.rendering.test.tsx | 411 ++++------------ .../TaskDetailModal.resolved-columns.test.tsx | 1 - ...etailModal.responsive-and-dependencies.test.tsx | 88 +--- .../__tests__/TaskDetailModal.spec-lock.test.tsx | 3 - .../__tests__/TaskDetailModal.summary-tab.test.tsx | 15 - .../TaskDetailModal.tab-persistence.test.tsx | 1 - .../TaskDetailModal.tab-relocation.test.tsx | 1 - .../TaskDetailModal.task-activity-chat.test.tsx | 8 - .../components/__tests__/TaskDetailModal.test.tsx | 37 -- .../TaskDetailModal.worktree-terminal.test.tsx | 4 - .../__tests__/core-modals-mobile.test.tsx | 15 +- .../task-move-affordance-removed.test.tsx | 526 +++++++++++++++++++++ .../__tests__/workflow-resolved-columns.test.tsx | 169 +------ .../app/components/dashboard/MainContent.tsx | 1 - .../app/components/useRightDockController.tsx | 4 +- .../app/task-modal-touch-resize-e2e-fixture.tsx | 3 +- packages/i18n/locales/en/app.json | 44 +- packages/i18n/locales/es/app.json | 44 +- packages/i18n/locales/fr/app.json | 44 +- packages/i18n/locales/ko/app.json | 44 +- packages/i18n/locales/pt-BR/app.json | 44 +- packages/i18n/locales/zh-CN/app.json | 44 +- packages/i18n/locales/zh-TW/app.json | 44 +- packages/i18n/src/resources.d.ts | 42 +- 63 files changed, 884 insertions(+), 2628 deletions(-) Fusion-Task-Id: FN-198 Fusion-Task-Lineage: 21198580-6146-47e7-ae91-6a878e2edd3e Co-authored-by: Fusion <noreply@runfusion.ai> |
||
|
|
017ebd5059 |
FN-201: park workspace review revisions for human resolution
Preserve multi-repository review findings and route unresolved REVISE outcomes to human review. - Qualify workspace findings by repository for durable scope and remediation tracking. - Preserve findings through graph outcomes and stabilize remediation signatures. - Park multi-repository review revisions for human resolution with regression coverage. Files changed: .changeset/fn-201-workspace-review-findings.md | 7 + docs/workflow-steps.md | 2 +- .../__tests__/workspace-review-findings.test.ts | 198 +++++++++++++++++ .../workspace-review-remediation-routing.test.ts | 235 +++++++++++++++++++++ .../executor/append-review-remediation-steps.ts | 116 +++++++++- .../request-pre-merge-optional-step-fix.ts | 7 +- .../engine/src/executor/run-graph-custom-node.ts | 54 +++-- .../src/executor/workspace-review-per-repo.ts | 46 +++- .../src/executor/workspace-review-remediation.ts | 7 +- 9 files changed, 638 insertions(+), 34 deletions(-) Fusion-Task-Id: FN-201 Fusion-Task-Lineage: cb83c4cc-b6fc-4e4c-bd3d-5c855bf2093b Co-authored-by: Fusion <noreply@runfusion.ai> |
||
|
|
c4e775eb43 |
FN-200: reuse unified chat model and keep thinking popover open
Reuse the unified Direct Chat model and thinking-level popover behavior across Task Chat surfaces. - Share the unified model selector and thinking control with Task Chat. - Keep the thinking popover open while selecting a model and align responsive styling. - Update documentation and regression coverage for the shared behavior. Files changed: .../fn-200-unified-chat-model-thinking-popover.md | 7 + docs/dashboard-guide.md | 8 +- .../app/components/ChatThinkingLevelControl.tsx | 172 +++++++++++++++------ packages/dashboard/app/components/ChatView.tsx | 1 + .../app/components/TaskPlannerChatTab.css | 50 ++---- .../app/components/TaskPlannerChatTab.tsx | 47 +++--- .../ChatThinkingLevelControl.portal.test.tsx | 5 +- .../__tests__/ChatThinkingLevelControl.test.tsx | 168 ++++++++++++++++++-- .../__tests__/ChatView.thinking-level.test.tsx | 4 + ...etailModal.responsive-and-dependencies.test.tsx | 9 +- .../__tests__/TaskPlannerChatTab.test.tsx | 86 ++++++++++- 11 files changed, 420 insertions(+), 137 deletions(-) Fusion-Task-Id: FN-200 Fusion-Task-Lineage: 16e9fbc7-89bf-4492-bed3-c30cdbf85393 Co-authored-by: Fusion <noreply@runfusion.ai> |
||
|
|
54403ffae9 |
FN-197: split task detail tabs and report file-scope blockers
Move task attachments, dependencies, blocking, and specification content into focused task-detail tabs while exposing overlapping files as a merge blocker. - Relocate task-detail definition sections into dedicated tabs with updated responsive styling and localization. - Add file-scope overlap reporting through dashboard routes and engine scheduling integrations. - Update affected tests, documentation, and the published-package changeset. Files changed: .changeset/fn-197-task-detail-tab-split.md | 7 + docs/dashboard-guide.md | 4 +- packages/dashboard/app/api/legacy.ts | 3 + packages/dashboard/app/api/tasks/tasks.ts | 14 + packages/dashboard/app/components/TaskDetailModal.css | 41 + packages/dashboard/app/components/TaskDetailModal.tsx | 1192 +++++++++++--------- packages/dashboard/app/components/__tests__/TaskDetailModal.attachments-and-tabs.test.tsx | 45 +- packages/dashboard/app/components/__tests__/TaskDetailModal.definition-actions.test.tsx | 12 +- packages/dashboard/app/components/__tests__/TaskDetailModal.github-tracking-header.test.tsx | 4 +- packages/dashboard/app/components/__tests__/TaskDetailModal.github-tracking-renamed-lanes.test.tsx | 2 +- packages/dashboard/app/components/__tests__/TaskDetailModal.github-tracking-stale.test.tsx | 16 +- packages/dashboard/app/components/__tests__/TaskDetailModal.gitlab-tracking.test.tsx | 4 +- packages/dashboard/app/components/__tests__/TaskDetailModal.inline-editing-and-integrations.test.tsx | 87 +- packages/dashboard/app/components/__tests__/TaskDetailModal.rendering.test.tsx | 20 +- packages/dashboard/app/components/__tests__/TaskDetailModal.responsive-and-dependencies.test.tsx | 30 +- packages/dashboard/app/components/__tests__/TaskDetailModal.spec-lock.test.tsx | 18 +- packages/dashboard/app/components/__tests__/TaskDetailModal.tab-persistence.test.tsx | 2 +- packages/dashboard/app/components/__tests__/TaskDetailModal.tab-relocation.test.tsx | 198 +++ packages/dashboard/app/components/__tests__/TaskDetailModal.test-helpers.ts | 2 + packages/dashboard/app/components/__tests__/TaskDetailModal.test.tsx | 42 +- packages/dashboard/src/__tests__/routes-task-overlap-blocker-report.test.ts | 61 + packages/dashboard/src/routes/register-task-workflow-routes.ts | 17 + packages/engine/src/__tests__/file-scope-overlap-report.test.ts | 67 ++ packages/engine/src/execution/file-scope-overlap-report.ts | 79 ++ packages/engine/src/index.ts | 13 + packages/engine/src/scheduler.ts | 48 + packages/i18n/locales/en/app.json | 18 + packages/i18n/locales/es/app.json | 19 + packages/i18n/locales/fr/app.json | 19 + packages/i18n/locales/ko/app.json | 19 + packages/i18n/locales/pt-BR/app.json | 19 + packages/i18n/locales/zh-CN/app.json | 19 + packages/i18n/locales/zh-TW/app.json | 19 + packages/i18n/src/resources.d.ts | 48 + 34 files changed, 1463 insertions(+), 745 deletions(-) Fusion-Task-Id: FN-197 Fusion-Task-Lineage: b328ef8c-9445-4154-b401-464d777e1e6e Co-authored-by: Fusion <noreply@runfusion.ai> |
||
|
|
1d08fcb30d |
FN-196: restore the New Task modal Start button
Restore the New Task modal's discoverable Start action with validated workflow routing. - Keep Start visible but disabled until a description is entered. - Support atomic Coding (Ideas) creation and eligible manual-intake promotion. - Track Start submission state and expand regression coverage and documentation. Files changed: .changeset/fn-196-new-task-modal-start.md | 7 + docs/dashboard-guide.md | 3 + packages/dashboard/app/components/NewTaskModal.tsx | 29 ++- packages/dashboard/app/components/TaskForm.tsx | 5 + .../app/components/__tests__/NewTaskModal.test.tsx | 219 +++++++++++++++++++-- .../app/components/__tests__/TaskForm.test.tsx | 21 +- 6 files changed, 263 insertions(+), 21 deletions(-) Fusion-Task-Id: FN-196 Fusion-Task-Lineage: f3b7ff36-8a41-4189-b1c2-6dcf1f18287a Co-authored-by: Fusion <noreply@runfusion.ai> |
||
|
|
808f4c603b |
FN-199: request detailed reasoning summaries
Request detailed reasoning bodies for Responses-family model traces while preserving provider fallback behavior.\n\n- Shape supported Responses payloads to request detailed reasoning summaries.\n- Retry once with the prior payload when detailed summaries are rejected.\n- Preserve titled trace bodies across dashboard surfaces and document the setting.\n\nFiles changed:\n .changeset/fn-199-reasoning-summary.md | 7 + docs/dashboard-guide.md | 2 +- docs/settings-reference.md | 4 + .../__tests__/ThinkingTrace.surfaces.test.tsx | 19 +++ .../components/__tests__/ThinkingTrace.test.tsx | 4 + .../src/__tests__/pi-reasoning-summary.test.ts | 150 +++++++++++++++++++++ .../__tests__/reasoning-summary-payload.test.ts | 102 ++++++++++++++ .../src/execution/reasoning-summary-payload.ts | 72 ++++++++++ packages/engine/src/pi.ts | 48 +++++++ 9 files changed, 407 insertions(+), 1 deletion(-) Fusion-Task-Id: FN-199 Fusion-Task-Lineage: a3179f05-e709-4bae-9f4a-5507810717d8 Co-authored-by: Fusion <noreply@runfusion.ai> |
||
|
|
c31c15a8cd |
FN-195: add plain-language delivery summaries
Add a plain-language What This Delivers summary to task planning prompts and present task details summary-first. - Add policy and prompt guidance for concise delivery summaries. - Collapse the task Definition tab into a summary-first plan presentation. - Add summary extraction, responsive styling, documentation, changeset, and regression coverage. Files changed: .changeset/fn-195-what-this-delivers.md | 7 + AGENTS.md | 1 + docs/contributing.md | 4 + docs/dashboard-guide.md | 2 +- packages/core/src/__tests__/agent-prompts.test.ts | 40 +++- .../__tests__/original-description-policy.test.ts | 28 +++ packages/core/src/agents/agent-prompts.ts | 26 ++- .../core/src/tasks/original-description-policy.ts | 6 + .../dashboard/app/components/TaskDetailModal.css | 16 ++ .../dashboard/app/components/TaskDetailModal.tsx | 56 +++++- .../TaskDetailModal.plan-summary.test.tsx | 177 +++++++++++++++++++++ .../app/utils/__tests__/taskPlanSummary.test.ts | 114 +++++++++++++ packages/dashboard/app/utils/taskPlanSummary.ts | 118 ++++++++++++++ packages/engine/src/__tests__/triage.test.ts | 18 ++- 14 files changed, 598 insertions(+), 15 deletions(-) Fusion-Task-Id: FN-195 Fusion-Task-Lineage: a816797b-6f16-4613-9403-665fd19b73ed Co-authored-by: Fusion <noreply@runfusion.ai> |
||
|
|
1936c78774 |
FN-194: prevent involuntary Kanban scroll from text selection
Prevent native text-selection autoscroll from competing with intentional Board panning while preserving editable controls. - Suppress selection across Board surfaces and restore it for editable descendants. - Document the scrolling behavior and add CSS and component coverage for responsive and grouped columns. - Add a patch changeset for the dashboard behavior fix. Files changed: .changeset/fn-194-board-text-selection.md | 7 + docs/dashboard-guide.md | 5 +- .../app/__tests__/board-text-selection.test.ts | 86 ++++++++++++ packages/dashboard/app/components/Board.css | 20 ++- packages/dashboard/app/components/Board.tsx | 5 + .../__tests__/Board.text-selection.test.tsx | 144 +++++++++++++++++++++ 6 files changed, 264 insertions(+), 3 deletions(-) Fusion-Task-Id: FN-194 Fusion-Task-Lineage: 65f46bed-76f1-4e5f-9dce-7e6ce26b5fa5 Co-authored-by: Fusion <noreply@runfusion.ai> |
||
|
|
e213b94f25 |
FN-193: open pop-out chats on their originating thread
Ensure chat pop-out windows open on the requested conversation and remain visibly separated without corrupting shared geometry. - Initialize pop-out chat state from the originating session. - Preserve per-window cascade offsets and shrink presentation geometry only when necessary. - Document the behavior and add focused regression coverage. Files changed: .../fn-193-chat-window-thread-and-cascade.md | 7 + docs/dashboard-guide.md | 8 +- packages/dashboard/app/App.tsx | 8 +- packages/dashboard/app/components/ChatView.tsx | 3 + .../dashboard/app/components/FloatingWindow.tsx | 102 ++++++++-- .../__tests__/ChatView.core-contracts.test.tsx | 17 +- .../__tests__/ChatView.new-chat-default.test.tsx | 39 +++- .../ChatView.pop-out-host-inventory.test.tsx | 12 ++ .../FloatingWindow.cascade-separation.test.tsx | 102 ++++++++++ .../components/__tests__/FloatingWindow.test.tsx | 127 +++++++++--- .../PoppedOutChatWindows.cascade.test.tsx | 28 +++ .../PoppedOutChatWindows.opens-on-thread.test.tsx | 221 +++++++++++++++++++++ .../PoppedOutChatWindows.stacking.test.tsx | 3 +- .../__tests__/PoppedOutChatWindows.test.tsx | 4 +- .../__tests__/useChat.initial-session.test.ts | 194 ++++++++++++++++++ packages/dashboard/app/hooks/useChat.ts | 12 +- 16 files changed, 823 insertions(+), 64 deletions(-) Fusion-Task-Id: FN-193 Fusion-Task-Lineage: d5dc3ea4-1041-4e3a-b665-378bdf32d669 Co-authored-by: Fusion <noreply@runfusion.ai> |
||
|
|
286dd0a7ae |
fix(workspace): block completion only on main-checkout commits
Uncommitted edits in a sub-repo main checkout no longer refuse `fn_task_done`. They emit `worktree:workspace-main-checkout-edit` with `outcome:"warned"`, `reason:"uncommitted-only"`, and their evidence enum instead. Two measured reasons. The workspace land path is the same `landOneRepo` / `landSquash` mechanic as single-repo, with `projectRootDir` set to the sub-repo main checkout, so a dirty tree there is already stashed (untracked included) -> fast-forwarded -> restored under `merger.allowDirtyLocalCheckoutSync`; refusing completion for a state the very next stage is built to absorb stops the board for nothing. And an in-scope status entry carried no timing evidence at all, so an operator editing the same feature was indistinguishable from an agent that skipped `fn_acquire_repo_worktree` -- while the refusal named an operator-only remedy in a message addressed to the agent, so the card could only loop. The dangerous cases keep their refusals: a task-attributed commit still returns `main_checkout_edit` (it would reach the shared branch unreviewed), and work that exists only in a main checkout still fails the acquired-worktree `no_commits` invariant that actually proves delivery. |
||
|
|
71b09e677d |
chore(release): v0.77.0-beta.9
Version bump via changesets. |
||
|
|
e95f9bcb30 |
chore: downgrade FN-171 changeset from major to minor
The auto-reload settings removal drops an opt-out toggle but does not break the published API surface, so a minor bump under category `feature` is the accurate release-note framing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
5769d5cd61 |
Merge remote-tracking branch 'origin/main'
# Conflicts: # packages/dashboard/app/components/DockTaskList.tsx |
||
|
|
7e537989dc |
Merge remote-tracking branch 'origin/main'
# Conflicts: # packages/core/src/__tests__/settings-defaults.test.ts # packages/dashboard/app/components/ChatView.tsx |
||
|
|
922e93cf2b |
FN-9213: Display reverted tasks as in-column labels
Keep reverted work in its workflow lane while preserving visible resolution controls across dashboard surfaces. - remove separate reverted sections from the board, list, and right dock - label reverted tasks in normal list groups and retain Delete and Revise actions - deduplicate reverted rows during optimistic/refetch overlap and cover desktop and mobile behavior - document the updated workflow and add a patch changeset Files changed: .changeset/fn-9213-reverted-label-not-column.md | 7 +++ docs/dashboard-guide.md | 4 +- packages/dashboard/app/components/Board.tsx | 61 ++++++++------------ packages/dashboard/app/components/Column.tsx | 9 ++- packages/dashboard/app/components/DockTaskList.tsx | 34 +++++------ packages/dashboard/app/components/ListView.css | 5 ++ packages/dashboard/app/components/ListView.tsx | 41 ++++++++----- packages/dashboard/app/components/TaskCard.css | 1 - .../app/components/__tests__/Board.test.tsx | 47 ++++++++++++--- .../app/components/__tests__/DockTaskList.test.tsx | 9 ++- .../app/components/__tests__/ListView.test.tsx | 67 ++++++++++++++++++++++ .../app/components/__tests__/TaskCard.test.tsx | 13 +++++ 12 files changed, 216 insertions(+), 82 deletions(-) Fusion-Task-Id: FN-9213 Fusion-Task-Lineage: 3626da0d-185b-4f4a-b59f-72cfdcd1b60c Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai> |
||
|
|
47d2890031 |
fix(cli): make the raw logs view actually reachable, and stop it shadowing [v]
Two defects in my own previous commit, both visible in one screenshot of the dashboard. IT DID NOTHING ON A WIDE TERMINAL. The escape was added inside the single-pane layout, which is the NARROW one. A wide terminal renders the grid layout instead — System / Stats / Utilities / Settings beside Logs — so the toggle flipped state and the screen did not change. The whole point is that no chrome survives a rectangular selection, and chrome is drawn by both layouts plus the header, so the escape has to happen before either is chosen. It now replaces the entire frame, above the layout choice, and a test pins that ordering plus the fact that only one place may render it. IT SHADOWED AN EXISTING LEGEND. The Utilities panel already advertises `[v] Auto-Kill Vitest` on the same screen. The two handlers are mutually exclusive at runtime — utility actions require the Utilities section, this branch requires log focus — so nothing actually clashed, but two different `[v]` legends visible at once is a UI anyone would misread. Raw mode is now Shift+V; lowercase v stays with auto-kill. pnpm lint 0 errors, CLI typecheck clean, dashboard-tui suites 138/138, test:gate green. |
||
|
|
61f26ca4e7 |
feat(cli): add a raw, selectable logs view to the TUI
Reported: the logs cannot be highlighted and copied without column artifacts. Two things cause that, and the code already named the second one: the Logs panel keeps a border, a title and a filter row, and sits between a header and a status bar — so a rectangular drag captures box-drawing characters and unrelated rows — while mouse reporting is deliberately ON for that panel to drive wheel scrolling, which swallows the click-drag before the terminal ever sees it. `[v]` now shows the log lines alone: no border, no title, no filter row, no header, no status bar, every line starting at column 0, and mouse reporting released so the terminal's own selection works. One trailing hint row stays, because a full-screen view with no visible way out is worse than one extra row. `[v]` or Esc returns, and the escape is ordered ahead of the expanded-entry escape so the two modes cannot fight. The rendered line shape matches the existing `[c]` single-line copy, so selecting with the mouse and copying with the keyboard produce the same text. `[c]` remains the path for one line; this is the path for a range, which no keyboard shortcut can express. pnpm lint 0 errors, CLI typecheck clean, dashboard-tui suites 134/134, test:gate green. |
||
|
|
4d2b42b068 |
fix(FN-WF): stop a successful merge aborting itself, and clear two merge-lane dead ends
Three root causes, all reported from one live multi-repository board, all ending as noise or
as a dead end an operator had to clear.
A SUCCESSFUL MERGE ABORTED ITSELF. The in-flight fence gives up ownership as soon as a card
leaves the resolved review lane — correct for a REVISE pulling the card back to
implementation, wrong for the move to the complete lane that the merge performs on success.
Measured: `all 2 sub-repo(s) landed — task → done` at 19:58:00.762, then
`Aborting active merge (left-review-lane-during-merge)`. Nothing was actually cancelled —
both repositories were already on main — but the primitive was torn down after the fact,
which is why one merge wrote `Workflow node merge requested merge` twice, 132ms apart. Fired
a few hundred milliseconds earlier it would abort a merge genuinely mid-flight. This is the
COLUMN half of what FN-184 fixed for the STATUS half, in the same file: "the fence revokes the
very merge it is guarding".
A DUPLICATE ENDED AS AN ERROR. MULT-024 was closed through the duplicate sentinel — no commits
expected, implementation must not proceed — and the merge boundary then demanded a pre-merge
node result it could not possibly have, terminalizing it with "operator action required". A
task that did exactly what was asked required human rescue. The structural proof asks "did the
planned implementation run"; it is meaningless for an authorized no-commit outcome. Exemption
narrowed by the shared `hasNonTerminalSteps` rule, so a card with unfinished work still faces
the full proof, and it waives nothing else: pre-merge approval and FN-8141's skipped-
verification guard still apply at the door.
MERGE CHECKS HAD NO RUNNER. The clean room exists to run the project's checks, and every
runner, linter and type-checker lives in devDependencies — but the install inherited an
ambient NODE_ENV=production and skipped them all. The executor said so in its own words
("the environment omitted devDependencies") and repaired itself; the clean room did not, and
its reviewer approved a merge whose tests could not run. The project already neutralizes this
for its own tests in scripts/test-changed.mjs; the lesson never reached the lane that
provisions checkouts.
pnpm lint 0 errors, test:gate green, engine typecheck clean, pipeline-smoke 93/93, and 328
tests green across the suites covering every defect reported today.
|
||
|
|
a6a3e5fada |
fix(dashboard): rescue an empty Feed when the card opens straight onto it
Reported from a live board: a task that is done, and demonstrably has journal entries, shows "(no activity)". Two mechanisms, each harmless alone. `stripTaskListHeavyFields` empties `log` and KEEPS every other field, `prompt` included. So an SSE task:updated payload for a task with a spec arrives with `prompt` present and `log: []`. The detail mount effect treats `"prompt" in task` as proof the prop is a complete TaskDetail and returns WITHOUT requesting the detail. `prompt` and `log` are stripped by different paths, so that proxy is false for exactly the payload above — the card adopts a log-less snapshot as if it were complete. The only rescue, refreshEmptyActivityFeed, was bound to a segment CHANGE. A card that OPENS on Feed — `initialTab: "logs"`, which is how a deep link and the board's activity affordance land — never changes segment, so "(no activity)" was permanent for that visit. The rescue now runs whenever an empty Feed is visible. Its existing emptiness guard is what keeps this cheap: a populated feed still costs no request, and a genuinely empty task asks once, because the callback identity is stable while it stays empty. Three regression tests: the stripped-snapshot open (fails without the fix), an honestly empty journal that must report empty without spinning, and a prop-carried journal that must render without re-requesting. One existing FN-8779 fixture gained a default mock implementation because it resets the mock and queues only one — its subject is Feed layout, not request counts, and no assertion was relaxed. pnpm lint 0 errors, dashboard typecheck clean, TaskDetail suites 758/758, test:gate green. |
||
|
|
c501ec9617 |
fix(dashboard): let Start work on a copied Ideas workflow, not just the built-in one
Start resolved its create-time Planning lane from the literal builtin:coding-ideas id, so a duplicated Ideas workflow fell through to a promotion that skipped the Planning hold lane and targeted the WIP lane. Column adjacency permits intake -> hold | archived only, so that move was rejected and the card stayed parked in Ideas. Resolve the lane from traits (first declared hold column immediately after a manual intake, mirroring resolveWorkflowIntakeFacts), and promote exactly one legal forward step when no atomic lane can be proven. |
||
|
|
cdef6ad7e8 |
fix(core): select the stale no-op merge case by its condition, not by a sentence
`merge-confirmed-finalize` carves out one case: a no-op merge confirmation with no landed commit is not proof the work was done, so when the steps are still unfinished the run must fall through to stale-merge cleanup and reverification instead of being consumed there. It selected that case by comparing the blocker reason with `===` against the exact string "task has incomplete steps". The merge-authority work then made refusals more informative, so a card in an error state reports `task is marked 'failed': … task has incomplete steps`. Same meaning, different sentence — and the comparison stopped matching, silently. A filter pinned to "subject is exactly Invoice" once invoices began arriving as "Invoice — March 2026". Nothing in the merge gate said so, because the test guarding this case lives in a file the gate does not run. It has been red on main since that lane landed. `hasNonTerminalSteps` states the rule the message describes and is defined from the same `NON_TERMINAL_STEP_STATUSES` set as `getTaskMergeBlocker`, so the two cannot drift. A blocker message is written for an operator and will be reworded again; the condition underneath it is what callers actually mean. The new core test pins them apart deliberately: it asserts the sentences DIFFER between a plain card and a failed one while the rule answers the same, and that the rule agrees with the door for every step status. A future prefix cannot re-break this quietly. pnpm lint 0 errors, test:gate green, core + engine typecheck clean, pipeline-smoke 93/93, ce-workflow-step-executor 53/53 (was 52/53 on main). |
||
|
|
7b9f839252 |
fix(FN-WF): ask for the verdict in a way that covers the case that broke it
An audit of how the verdict is REQUESTED, prompted by a reviewer that answered in prose. Prompt text only; no parser, type or lifecycle change. This block is the last thing in a review step's system prompt, so what it says last carries the most weight. It said this: "Backward compat fallback: if JSON is unavailable, you may still begin output with REQUEST REVISION". The closing words of the entire prompt granted permission to skip the required format, on a false premise — emitting JSON is never unavailable. An imperative followed by a dispensation is a preference. The degraded path still exists in the parser, but is now described as degraded rather than as an alternative, and no longer occupies the final line. It also forbade markdown fences while the parser scans fenced blocks FIRST, so "compliant" was narrower than "parseable" for no benefit, against a habit most models have. And the one that actually explains the incident: it offered APPROVE, APPROVE_WITH_NOTES and REVISE, with no legal way to say "I cannot see the change I was asked to review". On the measured multi-repo card the reviewer was told no files had changed, found nothing, and none of the three values described its situation — approving would have been a lie. So it wrote prose, which the gate then swallowed. The model did not go off-format by accident; it was asked to choose from a list that did not contain its answer. That case is now explicitly mapped onto REVISE with the search stated in notes. A dedicated UNAVAILABLE member would model it better, and was deliberately NOT added: `WorkflowStepVerdict` has no such value, and introducing one reaches the parser, recorded step results, merge admission and the dashboard — out of proportion to a prompt repair, and outside what "no negative impact" permits. Impact checked before and after: the only test touching this text asserts the `## Feedback Format` heading, which is preserved. The one failure in that file reproduces identically with these changes stashed — it belongs to the merge-authority lane of 2026-08-23. pnpm lint 0 errors, test:gate green, engine typecheck clean, pipeline-smoke 93/93. |
||
|
|
828be7648b |
fix(FN-WF): stop the journal announcing aborts that never happened, and assert it
The operator journal is a deliverable. Nothing asserted it, and three defects lived there. ABORT BREADCRUMB. `awaitAbortInFlightTaskWork` wrote `Pause abort marked` before inspecting any surface, so a card with no session still announced an interruption: every newly created task logged `provenance=hard-cancel` a second after creation, because creation moves the card out of the planning lane and that move is user-sourced. Nothing was interrupted and the operator withdrew nothing — false on both counts, and the second time this label has lied. The in-memory marker is still claimed synchronously, before any await, because the graph-failure classifiers depend on it; only the operator-facing line waits for evidence. DUPLICATE APPROVAL. Landing requires TWO consecutive clean approvals of the same candidate. Both wrote the identical sentence with the identical SHA, so a safety feature read as a duplicated invocation and was reported as an anomaly. The line now carries its pass number. DEAD RECOVERY. That same line is a contract: SelfHealingManager parses it with `/AI merge review \(pass \d+\): approved …/` to recover approved SHAs. No emitter ever wrote the parenthetical, so the parser matched nothing, `hasApprovedAiMergeReview` always answered false, and the recovery it guards could not run. Two sides individually reasonable, coupled through a log line nobody compared — the same shape as every other defect in this series. Emitter and parser now agree, and a test pins them against each other so a one-sided edit fails instead of silently killing the path again. COVERAGE. New pipeline-smoke scenario S20 drives a task to merge on all three coding built-ins and asserts the journal an operator actually reads: no abort claimed on an uninterrupted card, no line written twice in a row, no approval whose own text says it verified nothing. It reproduced the duplicate deterministically on its first run, which is the point — every anomaly reported this week was plainly visible in that journal and invisible to this lane. pnpm lint 0 errors, test:gate green, engine typecheck clean, pipeline-smoke 93/93. |
||
|
|
ca624f0584 |
fix(FN-WF): a blocking gate must not approve without a usable verdict
Restores FN-6582's rule, which a later operator request had relaxed — deleting its test
along with it.
Operator decision, now carrying the reason the first reversal lacked: the only legitimate
reason to stop a task is an LLM problem; everything else is fixed at the source, or the AI
is made unable to return anything but what is expected — and if it does anyway, restart
cleanly.
Restarting cleanly already happens, twice, inside executeWorkflowStep: a malformed primary
retries on the fallback model, or self-retries once on the primary when no fallback is
configured. So `malformed` reaching this decision does not mean "one fumbled response" — it
means the reviewer failed to return a usable verdict across every attempt. That IS the
LLM-class condition an operator accepts as a legitimate stop. What it must never mean is
approval.
Measured: a reviewer reported in prose that the deliverables were absent, carried no verdict
JSON, and the gate recorded success. Unreviewed work merged on a rejection nobody could see.
A prose classifier cannot close this — that text held no rejection marker at all ("revise",
"reject", "must fix" all absent) because it was a factual statement of absence. Only the
ABSENCE of a verdict is detectable, so absence must not approve.
Advisory gates keep the relaxation: a step that was never allowed to hold a card does not
start holding one, which is where the original operator ask actually applies.
pnpm lint 0 errors, test:gate green, engine typecheck clean, pipeline-smoke 90/90.
|
||
|
|
24adc4bf40 |
fix(FN-WF): review each workspace repository against its own diff base
A workspace Code Review runs the review step once per SUB-REPOSITORY worktree. `executeWorkflowStep` captured the reviewer's scope with the singular `task.baseCommitSha` regardless, and that base does not resolve inside a sub-repository: `captureModifiedFiles` returned [] and the prompt told the reviewer "(no modified files detected for this task)". Measured on a real multi-repo card whose executor had COMMITTED in both repositories. The reviewer went looking, could not see the committed fixtures inside its own scope, and reported them as never delivered — a confident, factual rejection produced entirely by a wrong diff base. It then vanished, because prose carrying no verdict JSON is classified malformed and passes a blocking gate. Two defects in series: one manufactured a false rejection, the other swallowed it. This fixes the first. The per-repo base was already recorded and already used by the evidence capture in workspace-review-per-repo.ts; it simply never reached the reviewer. `diffBaseCommitSha` threads it through, and a singular task with no override still uses the task field, so the ordinary path is unchanged. pnpm lint 0 errors, test:gate green, engine typecheck clean, pipeline-smoke 90/90. |
||
|
|
b956a7c8eb |
fix(core): repair a renumbered migration whose ledger row outlived its column
Reported from a dev instance: `column "memory_focus" does not exist` on every chat-session read, so the task planner chat 500s and never opens — with a startup that reports success. A ledger row asserts "a migration with this NUMBER ran". That is not the same claim as "this COLUMN exists" once a migration has been renumbered, and this one was renumbered four times — 0059 -> 0060 -> 0061 -> 0065 -> 0066 — each time because an upstream batch claimed the sequence first. A database can therefore carry a row from one numbering while a different migration owned that number on the boot that recorded it. The applier trusts the ledger absolutely, skips the migration, and leaves a schema that does not match it. Nothing fails at startup; everything fails afterwards, because Drizzle's `select()` emits the binary's full column list and one missing column breaks every read of the table. The defence already existed one table over: `0047` task recommendations verifies its materialized column in addition to the marker and replays its idempotent SQL. The lesson had been learned and not generalized. Both migrations renumbered on this branch — 0066 memory focus and 0067 session contention wait state — now carry it, and both SQL files are `ADD COLUMN IF NOT EXISTS`, so a replay over a healthy schema costs nothing. Two PostgreSQL regression tests reproduce the drifted state exactly (marker present, column dropped) and prove the replay materializes the column and stays idempotent on a second pass. pnpm lint 0 errors, test:gate green, core typecheck clean, schema-applier 80/80 against a real PostgreSQL. |
||
|
|
1c26a4bf4b |
fix(dashboard): report the cause of a failed query, not the statement that failed
Reported from a task chat: a screenful of column names from `project.chat_sessions` and nothing about what broke. That message is, by construction, the useless half. Drizzle wraps a query failure in an error whose message is `Failed query: <the whole statement> params: …` and puts the real PostgresError — `column "x" does not exist`, `permission denied`, `connection terminated` — in `cause`. `rethrowAsApiError` read `error.message` alone, so the reason was dropped before it ever reached the operator. `startup-factory` already carried a private chain walker because field reports of exactly this shape were undiagnosable; the dashboard never got one. The walker is now shared (`describeErrorChain` for logs, `summarizeErrorForOperator` for operator surfaces). The inversion is keyed narrowly on the `Failed query:` wrapper, never on guessing which message reads better: an application-authored message is deliberate prose and still leads, so the API boundary contract and its 29 tests are unchanged. Only the machine-generated frame is demoted to truncated context behind its cause. This does not fix the underlying query failure — it makes it reportable. The next occurrence will name the column or condition that failed instead of the statement that contained it. pnpm lint 0 errors, test:gate green, core + dashboard typecheck clean, 7 new tests. |
||
|
|
caae574146 |
fix(FN-WF): settle the restart scenario, and ban prompts that instruct denied tools
S17 ELUCIDATED. `restartPostMergeFinalization` read the task once, immediately after
restarting the engine, and treated "recovery has not finished yet" as "recovery will never
finish" — falling through to `admitAndMerge`. That fallback cannot succeed BY CONSTRUCTION:
staging deliberately replaces the row's step results with a single PENDING code-review row
and its steps with a pending stale step, precisely so the merge-confirmed recovery path is
what finalizes it. So merge admission was correctly refused and the scenario failed with
"post-merge restart parked finalization".
The outcome therefore depended on whether recovery beat one read: green in isolation
(19/19 across 8 runs) and intermittently red under full-lane load. That is a property of how
fast the suite happens to run, not of the product — the same conclusion FN-WF already reached
for S05, recorded in
|
||
|
|
9a54fe362d |
fix(FN-WF): give the Documentation milestone a way to actually persist anything
It had no writer, and its prompt did not know that.
A workflow step running `toolMode: "readonly"` is limited to read/grep/find/ls, fn_web_fetch
and a few read-only task reads; `fn_task_create` is explicitly DENIED there. The prompt asked
for four tool calls — fn_task_done(summary=…), fn_task_document_write, fn_artifact_register,
and creating follow-up tasks. It could make none of them. Every run produced a well-formed
report and persisted NOTHING. And because this milestone replaced `completion-summary`, which
used the working contract, cards quietly lost their agent-authored summary and fell back to
the deterministic backfill.
This is the same failure the reviewer prompt was fixed for — a session instructed to do what
its tool policy forbids — on a node nobody re-checked.
Both durable outputs now travel by PROJECTION, the only channel a writer-less node has.
`summaryTarget: "task"` persists its prose as the card summary. New
`recommendationsTarget: "task"` reads a trailing {"recommendations":[…]} payload, normalizes
it through the SAME rules the store boundary enforces (relocated to
tasks/recommendation-validation.ts so a second producer cannot drift from a copied regex),
and projects it to task.recommendations — the Recommendations tab, where an OPERATOR turns a
proposal into a task. An in-review agent proposes; it never creates board rows. Normalization
drops bad entries rather than throwing: a stray character in a suggestion must not wedge a
card whose code is already approved.
`summaryTarget` also removes this node's verdict requirement, so a reporter can no longer emit
the REVISE that held the merge door and bounced the card with nothing to do.
The guard that should have caught all of this asserted a PROMPT STRING —
`prompt.includes("fn_task_done(summary=")` — as proof a summary gets written. It was green
throughout. It now asserts the projection contract, including inside optional-group templates,
because the executing node of a group is its template child.
pnpm lint 0 errors, test:gate green, core + engine typecheck clean, pipeline-smoke 90/90.
|
||
|
|
56ee1622df |
fix(FN-WF): make Documentation a reporter, and refuse every bounce with no work
Observed on a live card (mult-021), where the log tells the whole story: Documentation returned an advisory REVISE asking for implementation work, the card was "moved back to in-progress for remediation", and 467ms later Code Review started again. No step was ever created, no executor session ran, and the demand was never implemented — the card merged when the second Documentation pass happened to pass. Two separate defects produced that. FIRST, the reporter could hold the merge. An advisory REVISE records `advisory_failure`, and `resolveRequiredPreMergeStepIds` included the Documentation group, so `evaluatePreMergeApprovals` read it as "not-approved". `gateMode: "advisory"` only stops the node blocking traversal; it says nothing to the merge door. SECOND, the reporter could bounce. `requestPreMergeOptionalStepFix` accepts `advisory_failure`, and under this workflow's named-remediation policy the resulting `sendTaskBackForFix` reopens NOTHING. With no pending step the foreach answered `already-expanded` and the walk replayed the review lane over an unchanged tree. The budget was 1/10, so it could have burned ten rounds of two model calls each. New opt-in `reportingOnly` on an optional group states the contract once — no approval to withhold, no remediation to request — and both doors read it. It is set only on Documentation, so advisory gates that DO own remediation (browser verification) keep their behaviour exactly. Plus the general invariant that would have caught both: under `stepReopenPolicy: "none"`, a bounce that appended no named steps is refused and logged on the card. Only the gates that can APPEND work may send a card back. Code Review REVISE and the deterministic verification failure still produce named fix steps — unchanged, still covered. pnpm lint 0 errors, test:gate green, core + engine typecheck clean, pipeline-smoke 90/90. |
||
|
|
8328b458b0 |
fix(FN-WF): prove fix steps reach the card, and clear the V2 rework's leftovers
FIX STEPS, asserted on `task.steps` rather than on a spy. A failing FINAL verification and a Code Review REVISE each append pending named steps carrying their gate provenance, and the card is re-dispatched to run them; completed implementation steps stay done, because remediation appends and never reopens. A review failure with NO REVISE verdict appends nothing — a transport error must not manufacture work. And no node id other than those two gates can reach the appender, which is what keeps a red test INSIDE a step the step's own problem: the executor fixes it there instead of littering the checklist. The new tests drive the real routing seam and the real appender against the real built-in registry — an injected IR is resolved away by workflow id and would have proved nothing. CATALOG. `builtin-workflows-lifecycle.test.ts` never received an EXPECTATIONS entry when V2 was registered, so its catalog-coverage assertion has been red on main since. The merge gate does not run that file, which is why it survived. Its trail is identical to builtin:coding-ideas by design: a read-only review lane changes what happens inside the working columns, not where the card goes. REGISTRY. The description still advertised "verify … summarize", steps that no longer exist, and the layout still positioned four deleted nodes plus drew Documentation to the LEFT of Code Review — so the editor rendered the review lane backwards against its own edges. Both now match the graph. AUDIT. `implementation-only-leakage` no longer flags `testing|verification`. That regex belonged to the revision where a review gate ran the tests; testing came back to the executor, so the planner emits that step on purpose and every V2 card was reporting leakage against its own intended plan. Documentation and delivery are still flagged. pnpm lint 0 errors, test:gate green, core + engine typecheck clean, 227 tests across the touched files. |
||
|
|
3cfb5119ea |
fix(FN-WF): stop V2 planning a Documentation step it already runs in review
Documentation is not a task step on this workflow.
Restoring the default planning prompt to bring `Testing & Verification` back also
restored `### Step {N}: Documentation & Delivery`, because the abandoned
`planning-implementation-only` seam stripped both in ONE anchored block, from the testing
heading to `## Documentation Requirements`. Nobody chose that; it was collateral.
The result was the same work done twice. The executor's step saved a delivery note,
registered artifacts and created follow-up tasks; the in-review Documentation milestone
then did the identical three tool calls again. Both wrote task document `docs`, so the
review pass silently overwrote the executor's.
`stripDocumentationDeliveryStep` removes ONLY the documentation block and deliberately
keeps `Testing & Verification`, which the executor owns and must keep planning. It is
applied to V2's own copy of the planning prompt, so `builtin:coding` and
`builtin:coding-ideas` keep the shared template byte-identical. If the base prompt is
reworded and the anchors stop matching, the strip degrades to an appended prohibition
rather than breaking planning at runtime.
Repository documentation survives as implementation work: the executor updates a doc its
own change made wrong, inside the step that made it, so Code Review sees it in the same
diff it approves. Whether a change warrants that is the executor's judgement, not a stage.
pnpm lint 0 errors, test:gate green, core typecheck clean, 2 new tests plus a shared-template
non-regression assertion.
|
||
|
|
0c3fe18347 |
fix(FN-WF): make a red verification create named fix steps, not an empty bounce
The FN-3345 deterministic verification gate runs testCommand/buildCommand after every planned step succeeds and before the in-review handoff. Both of its bounces went through `sendTaskBackForFix` regardless of the workflow's `stepReopenPolicy`. Under `none` — declared by `parse.implementationOnlySteps` + `preserveRemediationSteps`, selected today only by builtin:coding-ideas-v2 — that call reopens nothing, because send-task-back-for-fix.ts guards the reopen on `reopen-trailing`. So a card bounced back to implementation with ZERO pending steps, the foreach answered `already-expanded`, and it walked on to Code Review with the failing command unaddressed. The verification was measured, logged, and then silently discarded. The bounce shape now lives in executor/bounce-verification-failure.ts. `none` routes to `appendReviewRemediationSteps`, which derives one named step per file in the failing output, widens the PROMPT.md File Scope to those files, re-dispatches the executor, and parks for a human after three waves. `reopen-trailing` keeps its exact prior call, so builtin:coding and builtin:coding-ideas are byte-identical. This revives the `Verification` branch of appendReviewRemediationSteps, caller-less since the graph's `verification` node was removed — which is why the gap was invisible: the code was present, correct, and dead. pnpm lint 0 errors, test:gate green, engine typecheck clean, 5 new behavioural tests. |
||
|
|
b723c35fc9 |
feat(FN-WF): give testing back to the executor and the plan
Testing belongs to whoever can actually run it. That is the executor.
RESTORED — the planner emits "Testing & Verification" again. An earlier revision in
this series routed V2 planning through `planning-implementation-only`, whose contract
STRIPS that step region and replaces it with "Do NOT emit a Testing & Verification
step", on the theory that a review-column gate would run the checks instead.
Nothing ever did. The deterministic gate was not routed by its node kind and reported
PASS in ~46ms without executing anything; and once that was fixed, a review node runs
`toolMode: "readonly"`, where `bash` is denied and `fn_run_verification` is not in the
allowlist — so a reviewer cannot run lint, tests or build no matter what its prompt
says. Measured on real cards: 19s and 23s "reviews" that silently read the diff alone,
and a plan bounced for "implementation steps include testing and verification work
that must be handled as review-column gates" AFTER the gate it named was deleted. The
planner was forbidden from planning tests while nothing else ran them.
What was stripped is the mature contract: real automated tests only ("typechecks and
builds are NOT tests"), per-step test authoring, a final lint/tests/typecheck/build
pass ordered before delivery, an explicit duty to update tests that encode behaviour
the task changes, and standing up a test framework when the project has none. Plan
Review no longer rejects a plan for containing any of it.
CHANGED — Code Review judges the TESTS rather than claiming to run them. It rules on
four things: they exist for the behaviour that changed; they are real runner-executed
assertions; they assert BEHAVIOUR and never a comment or date stamp; and they cover
the invariant, not only the reported repro. Then it reviews the code for what tests
miss. Telling a session to do what its tool policy forbids invites the one failure
worse than a missing check — a fluent claim that the check passed.
DELETED — `builtin:review-gated-coding`, rather than left deprecated. It SHARED the
documentation-delivery node with V2, so every change made for V2 silently changed a
second workflow nobody was maintaining. Its own success path could never complete
anyway (`workspace-review-seal-required`).
Tests updated to the reversals they now describe, each naming the measurement that
reversed it. Deleting the workflow also cleared a pre-existing remediation-loop
failure.
pnpm lint 0 errors, test:gate, verify:fast, engine-pipeline-smoke 90/90, and three
consecutive full runs: 142.7s, 140.9s, 135.2s of the 175s budget.
|
||
|
|
e71ccb9d70 |
feat(FN-WF): show the in-review stage as a badge, not a step list
An in-review card now renders its stage through the running-gate badge alone — Code Review, then Documentation, then Merging — with no progress bar, no counter, and no expandable step list. This reverts the review-lane progress section added earlier in this series. That change was a correct fix for the complaint at the time (the review lane showed nothing at all), but with the lane reduced to two milestones in a fixed order the badge already answers "where is this card", and a list of two rows plus every finished implementation step is noise on a board. It also removes a defect for free rather than by repair. The list is built from `task.enabledWorkflowSteps`, which is FROZEN on the card at planning time, so a card planned before its workflow changed rendered a milestone that no longer exists as permanently `pending` — a ghost row that could never resolve. Removing the deleted `verification` group created exactly that on in-flight cards. No list, no ghost, and no reconciliation pass to write and maintain. Cost, stated rather than hidden: a NON-BLOCKING gate that failed is no longer visible from the board — Documentation cannot hold a card, so a failed delivery note now merges silently and must be read on the card itself. Blocking failures are unaffected: a Code Review REVISE moves the card back to in-progress, which is the most visible signal the board has. Tests updated to the new truth, not around it: the review-lane assertions now expect `.card-progress` and `.card-steps-list` to be absent. pnpm lint 0 errors, typecheck, test:gate, and 738 tests across every suite that touches card progress (TaskCard, ListView, taskProgress, board-mobile, live-ticker). The full dashboard suite's 22 failures are pre-existing and unrelated — ChatView, voice dictation, model menus, process supervision — with no card or progress test among them. |
||
|
|
8b64b88bfd |
feat(FN-WF): make the V2 review lane Code Review -> Documentation -> merge
One gate that can hold a card, one milestone that reports, then the merge.
REMOVED — the separate deterministic `verification` group. It duplicated the
executor's own verification, it showed a green badge on projects that had
configured no command, and it split merge evidence across two authorities that
could disagree. Code Review now runs lint/test/build itself, so exit codes still
decide and a single node owns the verdict. Its prompt is APPENDED to rather than
edited, leaving the shared reviewer used by builtin:coding and
builtin:coding-ideas exactly as it was.
The evidence rule is the point: the reviewer must quote each command with its exit
code and output tail, and a verdict with no execution evidence is invalid. A
reviewer free to assert "tests pass" in prose reproduces the false green a silently
passing gate produced mechanically — and the fluent version is harder to spot.
Absent commands are reported, never treated as failure: a project that never
configured verification has never been refused a merge on that basis.
REMOVED — `completion-summary` as its own milestone. Documentation writes the card
summary in the same pass as the delivery note. One model call, not two.
CHANGED — Documentation now runs AFTER the review, which is the ordering its own
author intended ("runs after passing verification and code review") and which the
review seal previously forbade. It is legal because it no longer writes the
repository: it is advisory, read-only, and records a Fusion-side delivery note,
artifacts, follow-ups and the summary. Repository documentation belongs to the
executor during implementation — a docs change is a code change, and writing it
after approval put it outside the diff the reviewer signed off.
It also cannot veto any more. As a blocking gate it bounced a task whose own plan
forbade implementing anything, and that card looped through the review lane every
five minutes indefinitely.
The seal invariant got STRONGER, not weaker: no node other than the reviewer itself
writes anywhere in the review lane, so nothing can change after an approval. The
test asserts exactly that, and names the reviewer exclusion rather than filtering it
away silently.
pnpm lint 0 errors, test:gate, verify:fast, engine-pipeline-smoke 90/90, and three
consecutive full runs: 137.4s, 141.0s, 144.6s of the 175s budget. The 3 remaining
core failures are pre-existing and reproduce without this diff.
|
||
|
|
28a8205645 |
fix(FN-WF): run the Verification gate instead of silently passing it
Your Verification step completed in 46ms and reported PASS without executing anything. It had never run. `GateNodeRunner` recognised exactly two executable shapes, `prompt` and `scriptName`. A gate carrying `workflowAction: "deterministic-verification"` matched neither and fell through to the method's closing `return success`. The code that runs testCommand/buildCommand was never reached — so the strictest gate in the review lane was decorative, and it was supplying the merge evidence a task is allowed to rely on. The existing unit tests were green throughout, because every one of them called `runDeterministicVerificationGate` directly. Testing a function proves the function; it does not prove the graph calls it. The new wiring suite asserts the routing itself and fails when the fix is removed — verified by removing it. DRY: `verification-gate.ts` re-derived the command list and re-ran the loop, a second implementation of a rule that `runExecutorDeterministicVerification` (FN-3345, run-implementation.ts) already owned. The two had already drifted — the copy treated "no command configured" as a hard failure while the original treats it as not-applicable. The gate now delegates, so timeouts, per-command logging, and settings precedence can only be fixed in one place. NOT in this change, and deliberately so: recording an unrunnable gate as `skipped` rather than `passed`. It is the right model and it was implemented end-to-end, but `pre-merge-approval` clears a `skipped` step only for an audited operator bypass, so it made every task on a project without a test command unmergeable — 25 of 90 smoke tests. Narrowing the acceptance left one unexplained failure (S09 sentinel, 120s timeout). Shipping that half-understood would trade a visible false green for an invisible merge deadlock. FN-189 owns it with the full evidence. pnpm lint 0 errors, test:gate, engine-pipeline-smoke 90/90, and three consecutive full runs: 150.9s, 152.8s, 154.5s of the 175s budget. |
||
|
|
ea869ff38f |
fix(FN-186): give pipeline-smoke tasks process-unique ids and drain execution on teardown
S05 now covers builtin:coding-ideas-v2. 5/5 consecutive full lanes. Two harness defects, both of which made a CORRECT engine refusal look like a flake. 1. Task ids collided. The serial lived on the harness instance and reset with it, so every test's first task was `FN-182-S05-1`. The engine's process-wide state — `executingTaskLock`, `activeSessionRegistry.pathsForTask`, worktree registrations — is keyed by task id, so a straggler from the previous test answered for the NEXT test's identically-named task and handed it a worktree under the PREVIOUS fixture. The serial now lives on the module. 2. Teardown forgot in-flight work instead of waiting for it. `ProjectEngine.stop()` clears timers but does not await an execution already inside `execute()`, and the harness then called `activeSessionRegistry.clear()` — which hides a live session rather than ending it. `dispose()` now drains `executingTaskLock` and the registry for its own task ids, bounded, and THROWS on expiry: a straggler that outlives the budget is a real defect, and a silent continue would restore the leak. Throughout this, the product was right. The executor detected the foreign worktree, refused it (`outside_worktrees_dir`), retried, exhausted its budget and failed visibly. That refusal is the desired behaviour and was never the bug — the harness was manufacturing the condition. Budget re-baselined 150s -> 175s for attributable growth: a 7th file (the remediation drive) and S05 on V2, one of the longest scenarios. Five runs at 140.1-148.4s left under 2s of headroom against the old ceiling, which is a flake waiting to happen. The standing rule is unchanged and now has three precedents: growth must be nameable, or it is a regression to fix rather than a budget to raise. pnpm lint 0 errors, test:gate, verify:fast, engine-pipeline-smoke 90/90, and five consecutive full runs: 146.0s, 142.6s, 148.5s, 144.8s, 146.3s of the 175s budget. |
||
|
|
b39d66c002 |
test(FN-WF): settle at the manual-merge hold instead of guessing from a snapshot
`driveToManualMergeHold` returned on the first turn that merely LOOKED parked
(`column === "in-review" && reviewPassed`). That is a snapshot, and a review-column
workflow invalidates it one turn later: a Code Review REVISE appends remediation
steps and sends the card back to in-progress, so the caller's merge then hit the
engine's correct refusal ("task is in 'in-progress', must be in 'in-review'").
The engine was right and the driver was wrong. Suppressing that refusal would have
reproduced FN-175 exactly, which is why the earlier attempts to swallow it were
reverted rather than kept.
It now returns immediately on an authoritative `manual-required` work item, and
otherwise keeps turning until the observable signature (column, status, step
statuses, review statuses) stops changing across consecutive turns. That is a
property of the graph rather than of how fast the suite happens to run — which is
why the scenario was intermittent only under full-lane load.
S05 is deliberately NOT extended to builtin:coding-ideas-v2 in this change. It
passes 22/22 when its file runs alone but fails in the full lane, and the evidence
says the cause is harness isolation, not the product: the task is handed a worktree
belonging to a DIFFERENT fixture (observed .../fusion-pipeline-smoke-eXXRLy/...,
expected .../fusion-pipeline-smoke-K2iaTm/...). The engine detects this and refuses
it — `outside_worktrees_dir`, retried, budget exhausted — which is the correct
behaviour. Shipping that as a red scenario would be shipping a known flake, so the
coverage waits for the isolation fix.
pnpm lint 0 errors, test:gate, verify:fast, and three consecutive full runs:
139.2s, 135.1s, 139.9s of the 150s budget.
|