The cli-agent executor feature added engine.getCliAgentRuntime() (called by
the CLI dashboard command at load) and bumped @fusion/core SCHEMA_VERSION
108→110. Update the stale test stubs:
- cli dashboard.test.ts: add getCliAgentRuntime() to the mock FnAgent class
(returns undefined; runDashboard handles the no-runtime path). Fixes 52
failures all rooted at `cwdEngine?.getCliAgentRuntime is not a function`.
- roadmap-store.test.ts: the roadmap store layers on core's Database, so its
schema version tracks core; assert 110 instead of 108.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- ci-test-shard: timing outputFile is now RELATIVE — one pnpm invocation
fans out to several packages whose vitests all received the same
absolute path, so every package overwrote the same timings file (last
writer wins). Each package now writes <pkgDir>/.timings/; discovery
(discoverWorkspaceTimingFiles) and the CI artifact globs scan the tree
- acp event-bridge-bounds: 20s timeout on the CPU-bound plan-flood test
(timed out at default 5s under loaded CI shard, passes in isolation)
- acp process-manager: port-4040-allowlist marker for its doc comments
(main-side; local guard flagged it after merging main)
This PR advanced @fusion/core's SCHEMA_VERSION 102 → 105 (migrations 103
workflows, 104 task_workflow_selection, 105 orphaned-selection cleanup) but
the "reaches current version after init/migrate" assertions across the core
test suite — and the roadmap plugin's mirror test — still hardcoded 102. The
dashboard build break was masking this: the test shards never ran until the
build was fixed, then all four failed on `expected 105 to be 102`.
Updated every getSchemaVersion()).toBe(102) current-version assertion to 105
(db, db-migrate, goals-schema, insight-store, mission-store, run-audit,
store-merge-queue, merge-request-record, task-documents) plus the roadmap
plugin. agent-log-migration already asserts against the imported SCHEMA_VERSION
constant (the robust pattern); central-db asserts its own version 13 and is
unaffected.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
handlePlan charged the budget but never checked the ceiling or set the
flag, so a plan-ONLY stream kept emitting after crossing the cap (caught by
both review bots). It now flags + truncates exactly like text/thinking.
Adds the plan-only flood regression test (185 total) and the category
frontmatter field to the new solutions doc.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Users can now watch everything the agent does while a CE stage works, steer
it mid-stage, and read the whole conversation as a proper chat surface.
Live output:
- New host capability: CreateInteractiveAiSessionOptions.onProgress — the
engine adapter streams thinking/text deltas + tool start/end markers from
the pi agent hooks (any plugin can use this).
- Orchestrator buffers per-session live activity (merged deltas, discrete
tool lines, capped), emits throttled progress events over SSE, and
GET /sessions/:id attaches it as liveActivity for the polling fallback.
- Routes detach turn execution: start/answer/resume return immediately
(status active) and clients converge via push/poll — the turn is watchable
instead of hidden inside a blocking POST.
- Turn timeout is now INACTIVITY-based: an actively-working long turn is
never killed; a quiet one interrupts with its working trace preserved.
- On settle the trace persists into history as a condensed record.
Steering:
- Stage protocol: responses may be a direct answer, {value, comment}
(answer + guidance), or {feedback} (guidance without answering); the
system prompt instructs agents to treat steering as first-class input.
- CeFlow: guidance textarea alongside selectable questions — attach to the
clicked answer, or "Send guidance" on its own.
Q&A UI:
- Transcript no longer hides control records: past questions/answers render
as chat bubbles (option ids → labels), steering turns marked, working
traces as collapsible "Agent work" blocks, completion marker.
- Live working pane (pulse + streaming thinking/tool lines) while a turn runs.
Tests: 130 plugin tests green (14 new: live buffer/flush ordering, inactivity
watchdog survives active work, detached convergence, steering payload shapes,
transcript rendering, live pane). Engine seam tests green; plugin/core/
engine/dashboard tsc clean. Core full suite OOMs locally (known orchestrator-
shell issue) — covered by CI shards.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The store/orchestrator were already multi-session (independent rows + live
handles per session); this surfaces it end to end:
- Sessions panel in the dashboard view: lists every session with stage,
status badge ("needs your input" for awaiting_input), and last activity;
stays visible while a flow is open so switching is one click. Closing a
flow returns to the overview without stopping the session.
- useCeSession.open(): adopt an existing session (pins its projectId for
answer/resume/poll); useCeSessions list hook with push-event refresh and
poll fallback while any session is mid-turn.
- DELETE /sessions/:id + orchestrator.discard(): dispose the live handle
before deleting the row (pipeline-link rows kept for task provenance);
Discard affordance on settled sessions.
- Tests: cross-session independence through one orchestrator, store delete,
route list/delete, hook open/list/remove/push/poll, view panel
open/switch/discard. 116 tests green; plugin + dashboard tsc clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Two valid P1s from PR review threads:
- event-bridge: plan output bypassed the per-turn output cap — entry size was
bounded but entry COUNT wasn't (1000 entries ~ 64MB through onThinking).
Plans are now suppressed once the cap flags, capped at MAX_PLAN_ENTRIES=100
with a truncation marker, bounded, and charged to the budget. +2 tests.
- control-handler: with pauseForApproval but no findApprovalByDedupeKey, a
human approval was silently discarded (unreadable status -> deny). HITL now
requires BOTH closures upfront and default-denies before creating a request,
so no approval is wasted and no pending record orphaned.
184 tests green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- Check pauseForApproval BEFORE createApprovalRequest so a gate with
create-but-no-pause default-denies without orphaning a pending approval
record (greptile P1).
- prompt-builder: whitespace-only prompt yields no text block (code/comment
mismatch) + regression test.
- onLoad logs arg count, not raw args (args can carry inline tokens).
- Document that engine-driven session resume (loadAcpSession) is deferred v1.
- Strengthen tests: eviction path observed end-to-end, id-normalization
asserted via differing raw forms, loadSession receives the normalized id.
Skipped with reasons (recorded in review thread reply): exports-to-dist,
README title (package name is correct), Surface Enumeration boilerplate,
heavy-lift streaming-read/path-jail rework, and two suggestions that would
weaken the default-deny floor. 182 tests green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The echo-agent fixture's prompt handler awaits a sessionUpdate write before
registering the cancellable hang; the cancel notification is dispatched
concurrently and could land first on loaded CI shards, no-op, and leave the
prompt hanging forever (5s test timeout on shard 2). The fixture now records a
pending cancel so prompt() resolves 'cancelled' immediately regardless of
arrival order. Test-fixture-only change; verified 5x locally.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Tier-2 code review fixes:
- P1 correctness: EventBridge per-turn state never reset — once the per-turn
output cap tripped, all later turns were silently suppressed and tool/accum
state bled across turns. Surface resetTurn() and call it per prompt turn.
- P1: plan_update read a non-existent .entries field (wrong SDK shape) and
wiped the displayed plan — now a documented no-op (full 'plan' is source of
truth).
- P1 security: write-path TOCTOU — open without O_TRUNC, re-validate realpath,
then truncate, so an intermediate-symlink-swapped escaped target is never
truncated before rejection.
- DoS: fs read stat-gates and bounded-reads oversized files instead of loading
them fully before the ceiling.
- Security: stderr redaction now spans chunk boundaries; secret deny-list adds
.git-credentials/*.p12/*.pfx/*.keystore/.pgpass/.htpasswd/etc.
- Reliability: cancelAcpSession bounded by a timeout so a blocked stdin can't
delay the registry SIGKILL.
- Maintainability: drop dead ACP_NOT_IMPLEMENTED export; type agentCapabilities
via the SDK AgentCapabilities; strengthen the S1 write-denial assertion.
+4 tests (181 total); typecheck + eslint clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Post-implementation /simplify cleanup. Registers the plan-specified
process.on('exit', killAllProcesses) safety hook in index.ts (was missing —
closes an orphan-subprocess gap on hard exit). Removes the unwired idle-timer
(engine StuckTaskDetector + dispose()/registry teardown is authoritative per
KTD4a) and its tests. Fixes a stale dispositionFor doc comment, removes a
redundant identifier re-normalization in the event bridge, and clarifies why
the FusionCategory type keeps git_write/task_agent_mutation. Behavior-
preserving; 177 tests green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Wires the ACP runtime plugin into the published CLI (RUNTIME_PLUGIN_IDS in
tsup.config) and the on-demand BUILTIN_PLUGINS catalog (experimental), matching
the untrusted-subprocess security posture. Adds the Risk S1 default-policy
safety: an acpAllowUnrestricted acknowledgement (default false) — without it, a
blanket allow on a sensitive category is escalated to approval rather than
auto-approved under the allow-all default policy, applied in both the permission
floor and fs write gating. Adds docs/acp-contract.md (launch/readiness +
failure taxonomy), a README with the AGENTS.md-required upstream evidence
(SDK repo/docs/release/integrity), a bundle-output test for the staged plugin,
and a @runfusion/fusion minor changeset. Package green at 179 tests.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
path-jail.ts is a real symlink-resolving confinement jail (NOT the
project-root-guard string check): realpath validation within realpath(cwd),
parent-realpath + final-component lstat for new files (rejects dangling/symlink
finals), O_NOFOLLOW open + re-validation for TOCTOU, NUL/escape rejection, and
a deny-list for secrets (.env/*.pem/*.key/.npmrc/.netrc/id_*/credentials) and
git internals. fs-capabilities.ts: read honors line/limit + a hard byte
ceiling; write is default-OFF, size-capped, hard-rejects .git/**, and routes
through the file_write_delete gate (reusing the U5 floor) — block/require-
approval gate the write, never free. Handlers registered only when the
capability is enabled, consistent with the advertised fs capability. +39
tests (173 total), incl. real symlink-escape and .git-write rejections.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The agent is untrusted input and the high inactivity ceiling (KTD4) does not
bound an actively-flooding agent. Adds sanitize.ts (strip ANSI/control
sequences, bound strings, bound identifiers — reject path separators/NUL so an
agent-supplied id can never reach a path). event-bridge.ts now caps per-turn
cumulative output (5M chars, truncate-and-flag once) and per-chunk size (64k),
sanitizes text/thinking/tool-title before callbacks (S7), and bounds the
toolCallId correlation map with FIFO eviction (S5). sessionId passed through
boundIdentifier before storage. +28 tests (134 total).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The security floor for session/request_permission. Classifies each tool
call's kind into a Fusion action category and reads the per-category
disposition from the live policy (never a preset shortcut — S1/KTD3a), so a
custom rule blocking command_execution is honored even under the default
unrestricted preset. Selects allow_once only, never allow_always (S2).
Unmappable/missing/other kind and missing gate/policy default-deny;
require-approval routes through the gate's HITL closures (createApprovalRequest
-> pauseForApproval -> re-read status) or default-denies when no approver
exists. requestPermission tracks in-flight requests and drains them cancelled
on teardown (KTD4a). Couples only to a local PermissionGate (no @fusion/engine
import, KTD3). +29 tests (106 total).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Maps ACP session/update notifications to AgentRuntime callbacks using the
authoritative SDK 0.24.0 vocabulary: agent_message_chunk->onText,
agent_thought_chunk->onThinking, tool_call->onToolStart, tool_call_update
(completed/failed)->onToolEnd correlated by toolCallId, plan as full
replacement. tool-mapping.ts derives display names + normalizes args.
createSession now passes a bridging client handler into connect() so
streamed updates reach the engine callbacks. +24 tests (77 total).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Implements the real AgentRuntime: createSession spawns + handshakes (U2)
then opens session/new (empty mcpServers, KTD5), persisting sessionId, cwd,
and the engine-provided actionGateContext (KTD3) plus the live connection
on the session. promptWithFallback builds ContentBlocks and drives one
prompt turn to its terminal stopReason. cancel/loadSession/resume helpers;
dispose does best-effort cancel then registry-authoritative teardown (KTD4a).
prompt-builder.ts builds text/image ContentBlock[]. 8 files / 53 tests green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Adds the connection layer: spawnAgent + self-cleaning process registry,
env allow-list (no inherited process.env, KTD6b), redacted stderr capture
(S8), and connect() establishing a ClientSideConnection over ndJsonStream
and completing the initialize handshake with explicit integer protocol-
version negotiation (KTD2) under a timeout. fs capabilities advertised only
when toggled (KTD6); teardown is registry-SIGKILL-authoritative (KTD4a).
probe.ts adds an async readiness probe with a failure taxonomy. Includes a
minimal runnable echo-agent fixture and 25 unit tests.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
New runtime plugin registering runtimeId 'acp', mirroring the
fusion-plugin-droid-runtime shape. Adds @agentclientprotocol/sdk@0.24.0
and an SDK smoke-import test that gates on the load-bearing exports
(ClientSideConnection, ndJsonStream, PROTOCOL_VERSION=1) so a breaking
SDK change surfaces at U1. Runtime adapter is a contract-conforming
skeleton (incl. describeModel); session driving lands in U2/U3.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
A new workspace-acyclicity invariant on main (run via the PR merge) flagged
compound-engineering -> @fusion/dashboard -> compound-engineering: the plugin is
listed in @fusion/dashboard's deps (for view loading) AND declared @fusion/dashboard
as a runtime dependency, which the cycle check (deps+devDeps) and the
'bundled plugins must not depend on host packages' check both reject.
The plugin's only @fusion/dashboard use is the type-only PluginDashboardViewContext
import. Drop the @fusion/dashboard dependency and resolve that type via an ambient
dashboard-interop.d.ts + tsconfig paths mapping (the fusion-plugin-dependency-graph
interop pattern). Breaks the cycle; the host passes the real context at runtime.
safeParse() previously only caught JSON syntax errors, so a semantically-wrong
but valid column ('null', '{}', a string) would rehydrate a non-array
conversationHistory that later crashed appendHistory's spread (and a bogus
currentQuestion). safeParse now takes a shape validator and falls back to []/null
on invalid shapes too. + regression test covering 'null'/'{}'.
Align the dependency graph plugin's dashboard interop declarations with the current dashboard contract.
- import ReactNode for plugin task card rendering support
- add DetailTaskTab, PluginToastType, and PluginTaskView type exports
- update PluginDashboardViewContext to require workflowSteps and the expanded openTaskDetail signature
- add optional renderTaskCard and addToast hooks to match dashboard expectations
Files changed:
.../src/dashboard-interop.d.ts | 15 ++++++++++++---
1 file changed, 12 insertions(+), 3 deletions(-)
Fusion-Task-Id: FN-5935
Fusion-Task-Lineage: 32313a96-1008-4de6-a7cb-e6bfbc385534
- session-store: rowToSession parses JSON columns via a safeParse helper, so a
corrupted currentQuestion/conversationHistory column degrades to null/[] instead
of throwing and crashing reads of an otherwise-valid row (+regression test)
- session-routes: validate the ?status= list filter against CE_SESSION_STATUSES
(asCeSessionStatus) instead of casting an arbitrary query string with 'as never'
- _harness: makeScriptedSession throws on an empty script rather than yielding
undefined, surfacing test mistakes loudly
- useCeSession tests: use fake timers (advanceTimersByTimeAsync) instead of real
setTimeout waits for deterministic poll-interval assertions
Skipped: 3 doc nits in src/skills/ce-*/references/** — those are pinned upstream
ce-* skill copies (KTD5 vendored snapshot), not this repo's content.
Replaces polling-only with a true server-push seam (any plugin benefits):
- core: createRouteContext accepts an emitEvent override (default still logs)
- dashboard: emitPluginCustomSseEvent forwards a plugin's ctx.emitEvent calls to
connected /api/events clients as a project-scoped 'plugin:custom' event; the
plugin route context's emitEvent is wired to it
- dashboard: PluginDashboardViewContext gains subscribePluginEvents so views
consume push via a host capability (no raw EventSource, no deep app import)
- CE view subscribes its session to push and refetches on each event; polling
stays as the fallback when push isn't wired or an event is missed
Also: skill-reachability test installs into a temp dir (no repo-dir writes).
Tests: dashboard sse +2 (plugin:custom relay + project scoping), plugin 99.
The session's owning store (and its live in-process handle) is selected per
request by projectId. start() sent projectId but answer/resume/getSession did
not, so any project-scoped session broke on the first answer (a different store
resolved → session not found / no live handle). The client now captures the
start projectId and reuses it on every subsequent call. Closes the multi-project
session-identity residual.
Closes the U2/U5 skill-discovery carry-forward so the plugin's interactive ce-*
sessions actually load the stage's bundled skill in a live agent (not just in
scripted-fake tests).
Root cause: createFnAgent built its DefaultResourceLoader without forwarding any
skill-discovery path, and the interactive seam options couldn't carry one. The
loader's skillsOverride only *filters* skills already discovered from cwd's
standard roots, so the plugin-local .fusion-ce-skills/<id>/SKILL.md was never
discoverable.
Fix (end-to-end):
- AgentOptions.additionalSkillPaths forwarded into DefaultResourceLoader
- CreateInteractiveAiSessionOptions gains requestedSkillNames + additionalSkillPaths
- the interactive engine adapter forwards them to createFnAgent (skills +
additionalSkillPaths)
- the orchestrator runs the session with cwd on the real project root and hands
it [stage.skillId] + the install root
Proven: a real DefaultResourceLoader with additionalSkillPaths discovers ce-plan
and filters out ce-work; the orchestrator passes the right id/path/cwd. Plugin 96,
engine 136, core 99 tests green.
- reconciler: a deleted current-stage board task no longer wedges the pipeline
in 'running' forever (terminality computed over existing tasks only; all-deleted
is a no-op, not a wedge)
- session-store: a human-slow awaiting_input session is no longer misclassified
stale (interval rubric applies only to in-flight active/launching turns)
- stage-registry: pipeline progression uses an explicit order ordinal instead of
registry insertion order, so out-of-order registration can't corrupt advancement
- orchestrator.answer(): validate questionId before mutating state, so a stale id
can't destroy the persisted currentQuestion recovery anchor
- orchestrator.resume(): rehydrate a live interactive session by replaying
persisted history (side effects suppressed) so a resumed session is actually
answerable instead of dead-ending; honest interrupted+error fallback when no
factory is available
95 tests (6 new regression tests, each confirmed failing pre-fix).
Quality cleanup across the 9-unit build (behavior-preserving, 89 tests green):
- extract createCeTaskWithLink so the work bridge and reconciler share one
provenance+link contract (prevents drift)
- discovery list scan probes readability via accessSync instead of reading and
discarding full file bytes
- makeError helper replaces ~7 duplicated CeArtifactError literals
- shared asString route helper; drop dead pipelineIds set; collapse a double
pipeline-state write and a redundant link re-query in advance
- resolveStageSkillCwd no longer takes params it ignores
Add settingsSchema (default session provider/model, enabled stages, sync
reconcile-on-hooks toggle, reconcile cadence hint) aligned across manifest.json
and the runtime manifest, with typed getters. Wire the consumed getters:
orchestrator passes defaultProvider/defaultModelId into sessions and rejects
disabled stages; hooks gate their reconcile drain on reconcileOnHooks. Add the
plugin README and docs/plugins/compound-engineering.md documenting the hub,
interactive sessions, work bridge, and the sync ownership model.
reconcileIntervalMinutes is exposed as an operator cadence hint but not yet
consumed (no host scheduler; reconcile is on-demand by design).
Add a ce_pipeline_state machine (currentStage/status) kept distinct from
board-task ownership (task column) — separate tables, no shared column, per
FN-5719. onTaskMoved/onTaskCompleted hooks do only an indexed lookup + enqueue
and return well under the 5s budget; advancement happens in an on-demand
reconciler sweep that re-derives correct pipeline state from board truth, so a
dropped hook event still converges (no tight poll loop). Outbound CE-flow
changes create the next-stage board task. Conflict policy: board authoritative
for task state, CE flow authoritative for artifact/pipeline content (inbound
read and outbound write target different rows, so they cannot contend).
Host-scheduler note: no host timer wired; sweeps run on hook-drain and on the
dashboard/route refresh surface (tracked for the host event-publish follow-up).
When the work stage completes with a derived task list, create Fusion tasks via
ctx.taskStore.createTask tagged CE-originated (sourceType workflow_step +
sourceMetadata marker) and record an authoritative ce_pipeline_links row per
task (back-reference lives in the link table, not task-row JSON, per FN-5719).
Tasks then run the normal lifecycle untouched. Zero derived tasks is a clean
no-op. The link store is intentionally minimal for U8 to extend with the
bidirectional pipeline-state machine.