This PR advanced @fusion/core's SCHEMA_VERSION 102 → 105 (migrations 103
workflows, 104 task_workflow_selection, 105 orphaned-selection cleanup) but
the "reaches current version after init/migrate" assertions across the core
test suite — and the roadmap plugin's mirror test — still hardcoded 102. The
dashboard build break was masking this: the test shards never ran until the
build was fixed, then all four failed on `expected 105 to be 102`.
Updated every getSchemaVersion()).toBe(102) current-version assertion to 105
(db, db-migrate, goals-schema, insight-store, mission-store, run-audit,
store-merge-queue, merge-request-record, task-documents) plus the roadmap
plugin. agent-log-migration already asserts against the imported SCHEMA_VERSION
constant (the robust pattern); central-db asserts its own version 13 and is
unaffected.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
handlePlan charged the budget but never checked the ceiling or set the
flag, so a plan-ONLY stream kept emitting after crossing the cap (caught by
both review bots). It now flags + truncates exactly like text/thinking.
Adds the plan-only flood regression test (185 total) and the category
frontmatter field to the new solutions doc.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Users can now watch everything the agent does while a CE stage works, steer
it mid-stage, and read the whole conversation as a proper chat surface.
Live output:
- New host capability: CreateInteractiveAiSessionOptions.onProgress — the
engine adapter streams thinking/text deltas + tool start/end markers from
the pi agent hooks (any plugin can use this).
- Orchestrator buffers per-session live activity (merged deltas, discrete
tool lines, capped), emits throttled progress events over SSE, and
GET /sessions/:id attaches it as liveActivity for the polling fallback.
- Routes detach turn execution: start/answer/resume return immediately
(status active) and clients converge via push/poll — the turn is watchable
instead of hidden inside a blocking POST.
- Turn timeout is now INACTIVITY-based: an actively-working long turn is
never killed; a quiet one interrupts with its working trace preserved.
- On settle the trace persists into history as a condensed record.
Steering:
- Stage protocol: responses may be a direct answer, {value, comment}
(answer + guidance), or {feedback} (guidance without answering); the
system prompt instructs agents to treat steering as first-class input.
- CeFlow: guidance textarea alongside selectable questions — attach to the
clicked answer, or "Send guidance" on its own.
Q&A UI:
- Transcript no longer hides control records: past questions/answers render
as chat bubbles (option ids → labels), steering turns marked, working
traces as collapsible "Agent work" blocks, completion marker.
- Live working pane (pulse + streaming thinking/tool lines) while a turn runs.
Tests: 130 plugin tests green (14 new: live buffer/flush ordering, inactivity
watchdog survives active work, detached convergence, steering payload shapes,
transcript rendering, live pane). Engine seam tests green; plugin/core/
engine/dashboard tsc clean. Core full suite OOMs locally (known orchestrator-
shell issue) — covered by CI shards.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The store/orchestrator were already multi-session (independent rows + live
handles per session); this surfaces it end to end:
- Sessions panel in the dashboard view: lists every session with stage,
status badge ("needs your input" for awaiting_input), and last activity;
stays visible while a flow is open so switching is one click. Closing a
flow returns to the overview without stopping the session.
- useCeSession.open(): adopt an existing session (pins its projectId for
answer/resume/poll); useCeSessions list hook with push-event refresh and
poll fallback while any session is mid-turn.
- DELETE /sessions/:id + orchestrator.discard(): dispose the live handle
before deleting the row (pipeline-link rows kept for task provenance);
Discard affordance on settled sessions.
- Tests: cross-session independence through one orchestrator, store delete,
route list/delete, hook open/list/remove/push/poll, view panel
open/switch/discard. 116 tests green; plugin + dashboard tsc clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Two valid P1s from PR review threads:
- event-bridge: plan output bypassed the per-turn output cap — entry size was
bounded but entry COUNT wasn't (1000 entries ~ 64MB through onThinking).
Plans are now suppressed once the cap flags, capped at MAX_PLAN_ENTRIES=100
with a truncation marker, bounded, and charged to the budget. +2 tests.
- control-handler: with pauseForApproval but no findApprovalByDedupeKey, a
human approval was silently discarded (unreadable status -> deny). HITL now
requires BOTH closures upfront and default-denies before creating a request,
so no approval is wasted and no pending record orphaned.
184 tests green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- Check pauseForApproval BEFORE createApprovalRequest so a gate with
create-but-no-pause default-denies without orphaning a pending approval
record (greptile P1).
- prompt-builder: whitespace-only prompt yields no text block (code/comment
mismatch) + regression test.
- onLoad logs arg count, not raw args (args can carry inline tokens).
- Document that engine-driven session resume (loadAcpSession) is deferred v1.
- Strengthen tests: eviction path observed end-to-end, id-normalization
asserted via differing raw forms, loadSession receives the normalized id.
Skipped with reasons (recorded in review thread reply): exports-to-dist,
README title (package name is correct), Surface Enumeration boilerplate,
heavy-lift streaming-read/path-jail rework, and two suggestions that would
weaken the default-deny floor. 182 tests green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The echo-agent fixture's prompt handler awaits a sessionUpdate write before
registering the cancellable hang; the cancel notification is dispatched
concurrently and could land first on loaded CI shards, no-op, and leave the
prompt hanging forever (5s test timeout on shard 2). The fixture now records a
pending cancel so prompt() resolves 'cancelled' immediately regardless of
arrival order. Test-fixture-only change; verified 5x locally.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Tier-2 code review fixes:
- P1 correctness: EventBridge per-turn state never reset — once the per-turn
output cap tripped, all later turns were silently suppressed and tool/accum
state bled across turns. Surface resetTurn() and call it per prompt turn.
- P1: plan_update read a non-existent .entries field (wrong SDK shape) and
wiped the displayed plan — now a documented no-op (full 'plan' is source of
truth).
- P1 security: write-path TOCTOU — open without O_TRUNC, re-validate realpath,
then truncate, so an intermediate-symlink-swapped escaped target is never
truncated before rejection.
- DoS: fs read stat-gates and bounded-reads oversized files instead of loading
them fully before the ceiling.
- Security: stderr redaction now spans chunk boundaries; secret deny-list adds
.git-credentials/*.p12/*.pfx/*.keystore/.pgpass/.htpasswd/etc.
- Reliability: cancelAcpSession bounded by a timeout so a blocked stdin can't
delay the registry SIGKILL.
- Maintainability: drop dead ACP_NOT_IMPLEMENTED export; type agentCapabilities
via the SDK AgentCapabilities; strengthen the S1 write-denial assertion.
+4 tests (181 total); typecheck + eslint clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Post-implementation /simplify cleanup. Registers the plan-specified
process.on('exit', killAllProcesses) safety hook in index.ts (was missing —
closes an orphan-subprocess gap on hard exit). Removes the unwired idle-timer
(engine StuckTaskDetector + dispose()/registry teardown is authoritative per
KTD4a) and its tests. Fixes a stale dispositionFor doc comment, removes a
redundant identifier re-normalization in the event bridge, and clarifies why
the FusionCategory type keeps git_write/task_agent_mutation. Behavior-
preserving; 177 tests green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Wires the ACP runtime plugin into the published CLI (RUNTIME_PLUGIN_IDS in
tsup.config) and the on-demand BUILTIN_PLUGINS catalog (experimental), matching
the untrusted-subprocess security posture. Adds the Risk S1 default-policy
safety: an acpAllowUnrestricted acknowledgement (default false) — without it, a
blanket allow on a sensitive category is escalated to approval rather than
auto-approved under the allow-all default policy, applied in both the permission
floor and fs write gating. Adds docs/acp-contract.md (launch/readiness +
failure taxonomy), a README with the AGENTS.md-required upstream evidence
(SDK repo/docs/release/integrity), a bundle-output test for the staged plugin,
and a @runfusion/fusion minor changeset. Package green at 179 tests.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
path-jail.ts is a real symlink-resolving confinement jail (NOT the
project-root-guard string check): realpath validation within realpath(cwd),
parent-realpath + final-component lstat for new files (rejects dangling/symlink
finals), O_NOFOLLOW open + re-validation for TOCTOU, NUL/escape rejection, and
a deny-list for secrets (.env/*.pem/*.key/.npmrc/.netrc/id_*/credentials) and
git internals. fs-capabilities.ts: read honors line/limit + a hard byte
ceiling; write is default-OFF, size-capped, hard-rejects .git/**, and routes
through the file_write_delete gate (reusing the U5 floor) — block/require-
approval gate the write, never free. Handlers registered only when the
capability is enabled, consistent with the advertised fs capability. +39
tests (173 total), incl. real symlink-escape and .git-write rejections.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The agent is untrusted input and the high inactivity ceiling (KTD4) does not
bound an actively-flooding agent. Adds sanitize.ts (strip ANSI/control
sequences, bound strings, bound identifiers — reject path separators/NUL so an
agent-supplied id can never reach a path). event-bridge.ts now caps per-turn
cumulative output (5M chars, truncate-and-flag once) and per-chunk size (64k),
sanitizes text/thinking/tool-title before callbacks (S7), and bounds the
toolCallId correlation map with FIFO eviction (S5). sessionId passed through
boundIdentifier before storage. +28 tests (134 total).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The security floor for session/request_permission. Classifies each tool
call's kind into a Fusion action category and reads the per-category
disposition from the live policy (never a preset shortcut — S1/KTD3a), so a
custom rule blocking command_execution is honored even under the default
unrestricted preset. Selects allow_once only, never allow_always (S2).
Unmappable/missing/other kind and missing gate/policy default-deny;
require-approval routes through the gate's HITL closures (createApprovalRequest
-> pauseForApproval -> re-read status) or default-denies when no approver
exists. requestPermission tracks in-flight requests and drains them cancelled
on teardown (KTD4a). Couples only to a local PermissionGate (no @fusion/engine
import, KTD3). +29 tests (106 total).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Maps ACP session/update notifications to AgentRuntime callbacks using the
authoritative SDK 0.24.0 vocabulary: agent_message_chunk->onText,
agent_thought_chunk->onThinking, tool_call->onToolStart, tool_call_update
(completed/failed)->onToolEnd correlated by toolCallId, plan as full
replacement. tool-mapping.ts derives display names + normalizes args.
createSession now passes a bridging client handler into connect() so
streamed updates reach the engine callbacks. +24 tests (77 total).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Implements the real AgentRuntime: createSession spawns + handshakes (U2)
then opens session/new (empty mcpServers, KTD5), persisting sessionId, cwd,
and the engine-provided actionGateContext (KTD3) plus the live connection
on the session. promptWithFallback builds ContentBlocks and drives one
prompt turn to its terminal stopReason. cancel/loadSession/resume helpers;
dispose does best-effort cancel then registry-authoritative teardown (KTD4a).
prompt-builder.ts builds text/image ContentBlock[]. 8 files / 53 tests green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Adds the connection layer: spawnAgent + self-cleaning process registry,
env allow-list (no inherited process.env, KTD6b), redacted stderr capture
(S8), and connect() establishing a ClientSideConnection over ndJsonStream
and completing the initialize handshake with explicit integer protocol-
version negotiation (KTD2) under a timeout. fs capabilities advertised only
when toggled (KTD6); teardown is registry-SIGKILL-authoritative (KTD4a).
probe.ts adds an async readiness probe with a failure taxonomy. Includes a
minimal runnable echo-agent fixture and 25 unit tests.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
New runtime plugin registering runtimeId 'acp', mirroring the
fusion-plugin-droid-runtime shape. Adds @agentclientprotocol/sdk@0.24.0
and an SDK smoke-import test that gates on the load-bearing exports
(ClientSideConnection, ndJsonStream, PROTOCOL_VERSION=1) so a breaking
SDK change surfaces at U1. Runtime adapter is a contract-conforming
skeleton (incl. describeModel); session driving lands in U2/U3.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
A new workspace-acyclicity invariant on main (run via the PR merge) flagged
compound-engineering -> @fusion/dashboard -> compound-engineering: the plugin is
listed in @fusion/dashboard's deps (for view loading) AND declared @fusion/dashboard
as a runtime dependency, which the cycle check (deps+devDeps) and the
'bundled plugins must not depend on host packages' check both reject.
The plugin's only @fusion/dashboard use is the type-only PluginDashboardViewContext
import. Drop the @fusion/dashboard dependency and resolve that type via an ambient
dashboard-interop.d.ts + tsconfig paths mapping (the fusion-plugin-dependency-graph
interop pattern). Breaks the cycle; the host passes the real context at runtime.
safeParse() previously only caught JSON syntax errors, so a semantically-wrong
but valid column ('null', '{}', a string) would rehydrate a non-array
conversationHistory that later crashed appendHistory's spread (and a bogus
currentQuestion). safeParse now takes a shape validator and falls back to []/null
on invalid shapes too. + regression test covering 'null'/'{}'.
Align the dependency graph plugin's dashboard interop declarations with the current dashboard contract.
- import ReactNode for plugin task card rendering support
- add DetailTaskTab, PluginToastType, and PluginTaskView type exports
- update PluginDashboardViewContext to require workflowSteps and the expanded openTaskDetail signature
- add optional renderTaskCard and addToast hooks to match dashboard expectations
Files changed:
.../src/dashboard-interop.d.ts | 15 ++++++++++++---
1 file changed, 12 insertions(+), 3 deletions(-)
Fusion-Task-Id: FN-5935
Fusion-Task-Lineage: 32313a96-1008-4de6-a7cb-e6bfbc385534
- session-store: rowToSession parses JSON columns via a safeParse helper, so a
corrupted currentQuestion/conversationHistory column degrades to null/[] instead
of throwing and crashing reads of an otherwise-valid row (+regression test)
- session-routes: validate the ?status= list filter against CE_SESSION_STATUSES
(asCeSessionStatus) instead of casting an arbitrary query string with 'as never'
- _harness: makeScriptedSession throws on an empty script rather than yielding
undefined, surfacing test mistakes loudly
- useCeSession tests: use fake timers (advanceTimersByTimeAsync) instead of real
setTimeout waits for deterministic poll-interval assertions
Skipped: 3 doc nits in src/skills/ce-*/references/** — those are pinned upstream
ce-* skill copies (KTD5 vendored snapshot), not this repo's content.
Replaces polling-only with a true server-push seam (any plugin benefits):
- core: createRouteContext accepts an emitEvent override (default still logs)
- dashboard: emitPluginCustomSseEvent forwards a plugin's ctx.emitEvent calls to
connected /api/events clients as a project-scoped 'plugin:custom' event; the
plugin route context's emitEvent is wired to it
- dashboard: PluginDashboardViewContext gains subscribePluginEvents so views
consume push via a host capability (no raw EventSource, no deep app import)
- CE view subscribes its session to push and refetches on each event; polling
stays as the fallback when push isn't wired or an event is missed
Also: skill-reachability test installs into a temp dir (no repo-dir writes).
Tests: dashboard sse +2 (plugin:custom relay + project scoping), plugin 99.
The session's owning store (and its live in-process handle) is selected per
request by projectId. start() sent projectId but answer/resume/getSession did
not, so any project-scoped session broke on the first answer (a different store
resolved → session not found / no live handle). The client now captures the
start projectId and reuses it on every subsequent call. Closes the multi-project
session-identity residual.
Closes the U2/U5 skill-discovery carry-forward so the plugin's interactive ce-*
sessions actually load the stage's bundled skill in a live agent (not just in
scripted-fake tests).
Root cause: createFnAgent built its DefaultResourceLoader without forwarding any
skill-discovery path, and the interactive seam options couldn't carry one. The
loader's skillsOverride only *filters* skills already discovered from cwd's
standard roots, so the plugin-local .fusion-ce-skills/<id>/SKILL.md was never
discoverable.
Fix (end-to-end):
- AgentOptions.additionalSkillPaths forwarded into DefaultResourceLoader
- CreateInteractiveAiSessionOptions gains requestedSkillNames + additionalSkillPaths
- the interactive engine adapter forwards them to createFnAgent (skills +
additionalSkillPaths)
- the orchestrator runs the session with cwd on the real project root and hands
it [stage.skillId] + the install root
Proven: a real DefaultResourceLoader with additionalSkillPaths discovers ce-plan
and filters out ce-work; the orchestrator passes the right id/path/cwd. Plugin 96,
engine 136, core 99 tests green.
- reconciler: a deleted current-stage board task no longer wedges the pipeline
in 'running' forever (terminality computed over existing tasks only; all-deleted
is a no-op, not a wedge)
- session-store: a human-slow awaiting_input session is no longer misclassified
stale (interval rubric applies only to in-flight active/launching turns)
- stage-registry: pipeline progression uses an explicit order ordinal instead of
registry insertion order, so out-of-order registration can't corrupt advancement
- orchestrator.answer(): validate questionId before mutating state, so a stale id
can't destroy the persisted currentQuestion recovery anchor
- orchestrator.resume(): rehydrate a live interactive session by replaying
persisted history (side effects suppressed) so a resumed session is actually
answerable instead of dead-ending; honest interrupted+error fallback when no
factory is available
95 tests (6 new regression tests, each confirmed failing pre-fix).
Quality cleanup across the 9-unit build (behavior-preserving, 89 tests green):
- extract createCeTaskWithLink so the work bridge and reconciler share one
provenance+link contract (prevents drift)
- discovery list scan probes readability via accessSync instead of reading and
discarding full file bytes
- makeError helper replaces ~7 duplicated CeArtifactError literals
- shared asString route helper; drop dead pipelineIds set; collapse a double
pipeline-state write and a redundant link re-query in advance
- resolveStageSkillCwd no longer takes params it ignores
Add settingsSchema (default session provider/model, enabled stages, sync
reconcile-on-hooks toggle, reconcile cadence hint) aligned across manifest.json
and the runtime manifest, with typed getters. Wire the consumed getters:
orchestrator passes defaultProvider/defaultModelId into sessions and rejects
disabled stages; hooks gate their reconcile drain on reconcileOnHooks. Add the
plugin README and docs/plugins/compound-engineering.md documenting the hub,
interactive sessions, work bridge, and the sync ownership model.
reconcileIntervalMinutes is exposed as an operator cadence hint but not yet
consumed (no host scheduler; reconcile is on-demand by design).
Add a ce_pipeline_state machine (currentStage/status) kept distinct from
board-task ownership (task column) — separate tables, no shared column, per
FN-5719. onTaskMoved/onTaskCompleted hooks do only an indexed lookup + enqueue
and return well under the 5s budget; advancement happens in an on-demand
reconciler sweep that re-derives correct pipeline state from board truth, so a
dropped hook event still converges (no tight poll loop). Outbound CE-flow
changes create the next-stage board task. Conflict policy: board authoritative
for task state, CE flow authoritative for artifact/pipeline content (inbound
read and outbound write target different rows, so they cannot contend).
Host-scheduler note: no host timer wired; sweeps run on hook-drain and on the
dashboard/route refresh surface (tracked for the host event-publish follow-up).
When the work stage completes with a derived task list, create Fusion tasks via
ctx.taskStore.createTask tagged CE-originated (sourceType workflow_step +
sourceMetadata marker) and record an authoritative ce_pipeline_links row per
task (back-reference lives in the link table, not task-row JSON, per FN-5719).
Tasks then run the normal lifecycle untouched. Zero derived tasks is a clean
no-op. The link store is intentionally minimal for U8 to extend with the
bidirectional pipeline-state machine.
Extend the stage registry with presentation metadata (icon/label/glob) so adding
a stage stays data-only. Add CeFlow renderer for text/single_select/multi_select/
confirm questions over the U5 polling session routes, with a visibly-marked chat
fallback (AE1) that still completes the stage. Wire the artifact-hub launcher to
list and start registered stages. Add the skill-interaction audit test producing
a measured rich-vs-chat coverage ratio (declared classification: 8/8=100% across
brainstorm/ideate/plan), failing on unclassified interactions.
Add allowlist-based artifact discovery over the conventional CE locations
(STRATEGY.md, docs/ideation, docs/brainstorms, docs/plans, docs/solutions,
CONCEPTS.md), grouped by stage, with error entries for unreadable artifacts and
no reads outside the allowlist. Add list/read/preview routes (self-contained
sandboxed HTML, data-section markers) and the primary CompoundEngineeringView
with explicit empty/partial/error states, viewport-gated fetch, and a short-TTL
cache. Bind the view via registerBundledPluginViews + dashboard workspace dep
(the real binding; manifest componentPath is cosmetic).
Add ce_sessions schema (onSchemaInit), CeSessionStore (via
ctx.taskStore.getDatabase()), and CeOrchestrator driving a stage's skill on the
U4 interactive seam: streams thinking/text, persists questions as
awaiting_input, writes the stage artifact on complete, and auto-saves +
emits an observable event on interrupt/error (never silent loss). Lifecycle:
launching/active/awaiting_input/completed/error/interrupted with resume and an
interval-relative staleness rubric. Start/answer/resume/get-state routes.
Streaming transport is client polling of the get-session-state route for v1:
plugin routes have no native SSE and the loader emitEvent is a logging stub, so
true server push needs a host publish-to-/api/events seam (tracked follow-up).
Skill discovery is cwd+prompt-based; forwarding install paths through the seam
is a tracked follow-up.
Bundle pinned copies of 7 CE pipeline-stage skills (strategy, ideate,
brainstorm, plan, work, code-review, compound) under src/skills/ and declare
them via PluginSkillContribution. Empirical finding: the skills contribution
alone does not make a SKILL.md resolvable in a session -- the engine ingests it
as a name only. So onLoad runs an idempotent, isolation-guarded physical install
into a plugin-local .fusion-ce-skills/ dir (never a global ~/.claude/skills),
which the engine skill-resolver can then discover. Proven against the real
loadSkills + resolveSessionSkills pipeline.
The Binary Release workflow stopped producing any GitHub Release assets
because every release had at least one failing build leg, and the
github-release job (needs: all four builds, no if:) was skipped whenever
any leg failed — suppressing even successfully-built platforms.
Root causes fixed:
- github-release: add `if: !cancelled()` + zero-artifact guard so a single
failing leg yields a partial release instead of none.
- setup-node-pnpm cache key: add runner.arch. runner.os is only
Linux/macOS/Windows, so arm64 runners restored x64 node_modules missing
native deps (@rollup/rollup-linux-arm64-gnu), crashing `pnpm build`.
- macOS CLI sign step: guard on APPLE_CERTIFICATE_BASE64 so unsigned
binaries still publish when certs are absent; add timeout-minutes: 30 to
build-binaries to avoid 24h runner hangs.
- dependency-graph plugin: replace unix cp/mkdir -p (failed on Windows
cmd.exe) with a cross-platform node copy script.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>