TaskCard inferred "unplanned" from steps.length === 0 while triage's
todo-discovery and the scheduler's dispatch filter both decide from
PROMPT.md seed-ness, so the badges disagreed with the engine in both
directions: a real spec that parsed to zero steps read as "Queued to
plan" while the scheduler already treated it as a WIP-slot candidate, and
a re-seeded card still carrying old steps read as "Ready" while triage
was about to plan it. Either way the badge sent operators to the wrong
cap.
Adds the shared isTaskAwaitingPlanning predicate (replan park, missing
spec, seed-vs-real content) used by both triage's discovery and a new
best-effort `awaitingPlanning` enrichment on GET /api/tasks. TaskCard
derives both badges from that one value — strict complements — and keeps
the step count only as a fallback for SSE payloads and older servers.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`findLatestByDedupeKey` read `targetContext` through the string-only `fromJson`.
In backend (PostgreSQL) mode that column is jsonb and Drizzle returns it ALREADY
PARSED, so the dedupe scan never matched: every gate retry minted a duplicate
approval request, and an approved grant could never be redeemed. The live
database shows the signature plainly — 17 approved requests, 0 completed.
Normalize both shapes in one place (`normalizeTargetContext`), applied at
`rowToRequest` and both dedupe scan sites, so a row resolves whether it arrives
as a JSON string (SQLite) or a parsed object (Postgres).
The regression test asserts shape-independence rather than the single reported
case: the same stored key must resolve in BOTH shapes, and must not match a
different key or an absent context in either. Mutation-checked — reverting the
scan sites fails exactly the parsed-object case.
Cherry-picked ahead of #2457, which carries the wider approval/permission
hardening pass, because this one is an active production defect on its own.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The client bundle aliases `@fusion/core` to the leaf `core/src/types.ts` to
keep Node-only dependencies out of the browser, so a package-root import of
`FUSION_CLIENT_HEADER`/`FUSION_DASHBOARD_UI_CLIENT` typechecked but failed
`vite build`:
"FUSION_CLIENT_HEADER" is not exported by "../core/src/types.ts"
Follow the documented pattern instead of widening the root alias: declare a
`./task-delete-attribution` subpath export, add the matching Vite alias ahead
of the broader `@fusion/core` key (Vite matches in order), register the module
in the browser-safe-core allowlist, and import the subpath from the client.
`task-delete-attribution.ts` has no imports at all, so it is a safe leaf.
`app/utils/detectContentLanguage.ts` already warned about exactly this trap;
the miss was mine for verifying with typecheck, lint and test:gate but not
`pnpm build`, which is one of the four checks CI blocks on.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Three related fixes, all originating from a `[api:error] Request failed`
log line showing a 500 on `GET /api/tasks/FN-8610/runtime-fallback`.
1. Missing/deleted tasks now return 404 instead of 500.
`getTaskImpl` signalled a miss with a bare `Error`, and route catches
only mapped errno `ENOENT` to 404 — a leftover from the file-backed
storage era. In Postgres mode nothing sets an errno code, so every
unknown/missing/soft-deleted/wrong-project read returned 500. Adds a
typed `TaskNotFoundError` (message byte-identical) plus a shared
`task-lookup-error` mapper applied across the task, session-diff,
git/GitHub, workflow and file-workspace route registrars. The same
bare throw existed on both archive-lifecycle delete paths, so
`DELETE /tasks/:id` was affected too.
2. 5xx logs now carry the origin stack.
`rethrowAsApiError` constructed a fresh `ApiError` from the message
and discarded the original, so the `FNXC:ApiErrorDiagnostics`
contract logged the rethrow site rather than the throw site — the
reported log entry had no stack at all. Threads `cause` through the
error factories and walks the chain (bounded, cycle-guarded).
3. Task deletions are attributable, and non-operator deletes notify.
`task:deleted` audit rows recorded `agentId: "system"` for every HTTP
delete, making an operator click indistinguishable from a script or
an agent; the calling agent's task id was accepted by the store and
then never persisted. Adds a `callerKind` union recorded in audit
metadata, tags every delete call site, and stamps a self-reported
`x-fusion-client` header from the dashboard client. When the caller
is `agent-tool` or `api-unattributed`, a best-effort notice is sent
to the operator mailbox; operator and engine deletes stay silent.
`x-fusion-client` is attribution, not authentication — anything can send
it. No delete-blocking, gating or permission logic is added here.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Replace empty backendMode stubs with real AsyncDataLayer paths: archive ID
reservation and isTaskArchivedAsync, orphaned task.json re-import, health
snapshots via checkPostgresHealth, settings/agent memory caches for sync
readers, async builtin prompt overrides, and self-healing audit/health
callers that previously used dead sync SQLite fallbacks.
Pin reconcileOrphanedTaskDirsImpl empty result and getDatabaseHealthImpl
always-healthy sentinel under backendMode so inventory category (e) stays
aligned with production self-healing and health call sites.
Drive real sync-reader stubs that empty-return under backendMode
(merge request, workflow selection/overrides/settings, run audit,
legacy step snapshot, settings/health) so category (e) of the
migration inventory stays pinned to shipped behavior.
Inventory analysis found exactly six read-only legacy openers; pin them in a
structural scan and assert incomplete archive guards stay SQLite-free in
backend mode so new production SQLite construction fails CI.
Groundwork for FN-8603's remaining ~14-minute wait. Liveness for a pending
review gate is judged purely by a 15-minute staleness floor because a lease
records WHO took it (`leaseOwner` = run id) but not WHERE, and under multi-node
every engine sees every other engine's leases. A fresh-but-unknown lease might
be running on a peer, so the floor was the only safe test -- and a lease left by
this node's own crashed process is indistinguishable from it.
Adds `WorkflowStepResult.leaseNodeId` plus an optional `LocalNodeLeaseIdentity`
argument to `classifyReviewLease`. One narrow new case: a lease stamped with the
caller's OWN node id whose `startedAt` predates the caller's process boot is
provably dead -- the process that could have owned it is gone -- so it
classifies as `reclaim` immediately rather than aging out. Deliberately narrow,
because widening it is a double-dispatch risk: absent (legacy) or peer node ids
keep the floor, and a lease taken by this process after boot is still adopted.
InProcessRuntime.start() resolves the local node id from CentralCore (fail-soft;
on error it stays undefined and floor-only semantics apply) and passes it to
SelfHealingManager. The graph executor stamps the field when deps.localNodeId is
set.
NOT YET WIRED, so this is inert in production and behavior is unchanged end to
end: `localNodeId` is not threaded from WorkflowGraphTaskRunner /
WorkflowTaskRuntime down into the executor deps, so no lease actually carries a
`leaseNodeId` yet. The reader is ready; the writer needs that pass-through
(WorkflowGraphTaskRunnerDeps gains the field, the runner forwards it, and the
runtime supplies this.localNodeId). Stopping here rather than half-threading it.
Verified: tsc clean on core and engine, pnpm lint clean, pnpm test:gate green
(299 + 70), core workflow-step-results suite green.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
FN-8603 sat in-review for ~36 minutes after an engine restart killed its Code
Review session 34 seconds in. It did recover on its own; the cost was latency,
not a terminal park.
Sweep ordering. reconcile-orphaned-pending-step-results PRODUCES the failed
results that recover-failed-pre-merge-steps CONSUMES, but in the periodic
maintenance list it ran ~15 entries after it. A step orphaned in cycle N was
therefore rewritten to failed only after recovery had already scanned, so
nothing re-ran it until cycle N+1. Moved it immediately before its consumer and
removed the now-duplicated later entry. Startup recovery already ordered the two
correctly.
Post-review fix budget. Default raised 3 -> 10 per operator request. Three
passes is below the observed convergence length for the gates this fallback
actually governs -- Browser Verification and custom optional gates -- since Plan
Review and Code Review already resolve to "unbounded" when unset, and exhausting
the budget parks the card for a human. The declaration default and five inline
`settings.maxPostReviewFixes ?? 3` call sites in executor.ts/self-healing.ts had
drifted into separate literals, so raising one alone would have left every
unset-settings path on the old value; they now share the exported
DEFAULT_MAX_POST_REVIEW_FIXES.
Not done, and why. Re-dispatching a restart-orphaned lease immediately at
startup is the change that would close the remaining ~14-minute wait, but it is
unsound as specified: liveness is judged by a 15-minute lease-staleness floor
because leases carry no node attribution, so treating a pre-boot lease as dead
would let one node orphan another node's genuinely running review. Needs a node
id on the lease record first. Left the floor intact.
Verified: tsc clean on core and engine, pnpm lint clean, pnpm test:gate green,
self-healing orphaned-pending-step-results and optional-step-revision suites
green.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The prompt fix stops planners writing the verdict in prose, but it relies on
every model reading one sentence correctly. This closes the hole underneath it.
When the finalize read finds no spec at all, the planner's streamed reply is
searched for a line that is exactly `DUPLICATE: FN-NNNN`. If found, the engine
writes the canonical marker file and continues — so marker parsing, keep/delete
resolution, and the sourceMetadata.nearDuplicateOf that renders the operator's
decision all run on the unchanged file contract rather than a second code path
that could drift from it.
Deliberately narrow. The marker must occupy a whole line, only the first counts,
and recovery is gated on the plan being genuinely absent — a planner that wrote
a real spec is never overridden by something it said in passing. The text tail
is bounded because the verdict lands in the closing summary, and it tees off
onText rather than reading AgentLogger, whose buffer is flushed on a timer.
Verified both directions: the tests fail without the recovery block, and the
"wrote a real spec while mentioning a marker" case keeps its spec.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Route process spawn/exit, verification success paths, MCP connect, skill info listings, createFnAgent/session bookkeeping, and executor dispatch chatter through FUSION_DEBUG so the operator log pane keeps real lifecycle outcomes.
The planning prompt said "do not write PROMPT.md" and, in the same breath,
"write DUPLICATE: {id} to the output file" — where the output file IS
PROMPT.md. A planner that took the first clause literally wrote no file and
reported the duplicate in prose.
The engine only ever reads the verdict from PROMPT.md's contents, so that
duplicate was invisible: the task failed deterministic validation as
"PROMPT.md file not found or empty", retried, terminalized to failed, emitted
a task-wedge mail, was recovered to todo by self-healing, and re-planned —
three full Opus planning cycles on FN-8600 before it was caught, with no
operator decision ever surfaced because sourceMetadata.nearDuplicateOf is only
set on the branch that parses the file.
Both prompt sites now say to write PROMPT.md with the marker as its entire
contents, and say why prose alone is not recorded.
Note the engine ordering is already correct — tryFinalizeExplicitDuplicateMarker
runs before validateGeneratedPrompt, and a worktree-local spec is recovered
first. Nothing to reorder; the file simply never existed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Recognise workflow-graph moves into in-review so gate entry no longer emits handoff-invariant violations, and split pause-abort provenance so engine teardowns are engine-abort instead of hard-cancel.
Follow-ups to running the pre-merge review gates in `in-review`. Each was
verified against the code before being fixed; one reported issue was
refuted and is noted below.
1. Symbol locks (packages/core/src/task-store/moves.ts)
FN-8306 made the lifecycle transition the symbol-lock RELEASE authority
but wrote no counterpart. That was harmless while a task only left WIP
at handoff/terminal; the gate crossing now releases the task's declared
symbols and the remediation node re-enters `in-progress` to edit the
same files in the same live worktree with its locks gone. Neither
acquire site (scheduler dispatch, claimDueWorkflowWorkItem) is on the
graph re-entry path. Adds a symmetric re-acquire on `!wip -> wip`.
Best-effort by design: a contended symbol logs and proceeds, which is
exactly the pre-fix posture, rather than parking the remediation behind
another holder and re-creating the stranding this change set removed.
2. Premature merge (packages/engine/src/self-healing.ts)
`recoverMergeableReviewTasks` was the only in-review sweep with no
liveness gate. The graph commits the column crossing at node entry and
writes the gate's pending lease two DB round trips later, and
`getTaskMergeBlocker` has no notion of "enabled but resultless", so in
that window the sweep could enqueue a merge with Code Review never run.
Filters `executingIds`, matching recoverGhostReviewTasks.
3. Orphan sweep (packages/engine/src/self-healing.ts)
The reported restart hazard is REFUTED: nothing re-attaches an in-review
graph run, so those leases are genuinely dead and marking them failed is
correct FN-8492 behavior. But the sweep also runs from periodic
maintenance in the same live process, where a tick between the lease
write and session registration could fail a gate that just started.
Honors a within-floor `classifyReviewLease`, matching the semantics Plan
Review already had. Cleanup of dead leases is delayed by the staleness
floor, not defeated. The audit event gains `needsOperatorBypass` for
`autoMerge:false` rows, which self-healing deliberately skips and only
fn_task_bypass_review can clear — previously indistinguishable from an
auto-recoverable rewrite.
4. Stall detection (packages/engine/src/planner-overseer.ts)
The `reviewer` and `merger` stages had no time-based check at all and
returned `progressing` unconditionally, so a hung gate produced no
signal however long it sat. Adds gate-anchored detection on both (a
plain in-review card with no reviewState resolves to `merger`, not
`reviewer`), keyed on the pending lease's own `startedAt` rather than
`columnMovedAt` so it cannot fire during a legitimate human merge-wait.
`cumulativeActiveMs` is documented, not changed: it now excludes gate
runtime, but adding the `timing` trait to `in-review` would count arbitrary
human merge-wait as active work — a worse distortion than the omission.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Code Review and Browser Verification now run with the card in `in-review`
instead of `in-progress`, so the board shows the card under review with the
running step as a badge (matching the Coding (Ideas) preset). Their paired
remediation nodes stay in `in-progress`, so a changes-requested verdict
visibly sends the card back to implementation.
The column move IS the badge switch: the dashboard badge was already
lane-gated on `column === "in-review"`. Applied to the shared stepwise
coding IR, so it is inherited by builtin:coding (the default),
builtin:stepwise-coding, builtin:brainstorming and builtin:coding-ideas;
builtin:legacy-coding keeps its historical placement.
Two consequences handled:
- Capacity: `in-review` has no `wip` trait, so the slot is released during
review and the remediation crossing back into `in-progress` can hit the
non-bypassable in-transaction capacity check. The column boundary now
PARKS the run on a `capacity-exhausted` rejection instead of failing it,
preserving the failed gate result and worktree so the next graph run
retries once a slot frees. Non-capacity rejections still propagate.
- Reopen clears: `applyReopenFieldClears` wiped `workflowStepResults` on
every in-review -> in-progress move, which the remediation crossing now
performs routinely. That destroyed the remediation input, made
`routeRetryableRemediationGraphFailureToPreMergeFix` and
`recoverFailedPreMergeWorkflowStep` silently no-op, and — worse — made
both `getTaskMergeBlocker` branches vacuously false, so a card could
return to `in-review` and be mergeable with its gate never re-run. Now
exempted for graph-owned in-review -> in-progress crossings only;
operator reopens, merge bounces and every -> todo/triage rebound still
clear, so the executor's documented bounce invariant is unchanged.
Adds regression coverage for both (there was previously none for the
reopen clear in either direction), and annotates the unreachable legacy
scheduler dispatch block rather than mirroring the fix into dead code.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Follow-up to 13a2b2a9d, from a multi-agent review of that commit. Three of its
claims did not hold.
1. fn_delegate_task bypassed the gate entirely (P0). It reaches the same
createAgentTask primitive, was registered unconditionally in both session
lanes, and validated only that the TARGET agent is non-ephemeral — never the
caller. Under Deny an ephemeral worker could enumerate agents and delegate
unlimited tasks. It is now withheld under Deny, and also under
upon_validation: delegation has no proposal channel, so leaving it available
would launder a create past the operator review that policy requires.
2. The widened dedupe window was capped at 5 minutes. The store query in
branch-and-pr-entities.ts carried its own independent `?? 60_000` /
`min(300_000, …)` pair, so widening only duplicate-guard.ts under-delivered
and made the new ceiling unreachable. Both sites now share
FINGERPRINT_WINDOW_DEFAULT_MS / FINGERPRINT_WINDOW_MAX_MS.
3. The pi-extension gate does not fire at all. pi's ExtensionContext carries no
agentId — the read is a speculative cast and only tests supply one, so every
real call short-circuits as a human caller. The fail-closed direction is kept
for the day an identity signal exists, but the limitation is now documented
instead of implied to be enforcement.
Also: the session prompt now states when creation is disabled and names
fn_task_log as the fallback (the base prompt still taught fn_task_create, which
is the same instruction/capability mismatch that fed the retry storm);
suppression emits an `agent:task-create-withheld` run-audit event; and the two
source-text ratchet tests are replaced with behavioral assertions on the tool
list the executor actually hands the model — verified to fail when the guard is
broken, which the string assertions did not.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Operator report: with project policy "Ephemeral agent follow-up tasks = Deny",
an executing agent filed ten follow-up tasks — five parallel fn_task_create
calls it reported as timed out, then five sequential retries.
Two defects:
1. Deny was advisory. fn_task_create was registered for every session and only
refused inside execute(), so the model still saw the tool, planned around it,
and retried it. The pi extension's isEphemeralCallerAgent also failed OPEN
whenever the caller id did not resolve to an agent row — which is the normal
shape of an ephemeral task-worker — so on that lane Deny was a no-op.
2. The deterministic content-fingerprint duplicate window was 60s, which only
covered concurrent in-flight creates. A retry two minutes later saw nothing
and filed a second task.
Fixes: isAgentTaskCreateToolAvailable() withholds the tool from ephemeral
sessions under Deny in both engine lanes (outer execution session, per-step
workflow session); isEphemeralCallerAgent fails closed on an unresolvable
caller id; the fingerprint window goes 60s -> 10m (clamp ceiling 5m -> 1h).
upon_validation keeps the tool, and permanent-agent and human/chat callers are
unaffected.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Add the `autoUpdateAndRestart` global setting (default off, Settings ->
General next to Release channel). When enabled, the dashboard host installs
available updates on the selected channel by itself and requests the
supervised in-place restart. Supervised hosts only: without a parent to
respawn, installing would leave a running process whose code no longer
matches its own install.
Fix two ways the restart affordance could silently do nothing:
- The supervisor now stamps FUSION_SUPERVISOR_PID and supervision is only
counted when that pid is the real parent. FUSION_RESTART_SUPERVISED is
inherited by every process Fusion spawns, so `fn dashboard` launched from
an agent terminal skipped its own supervisor while still advertising
restart support -- a restart request then killed it for good.
- Settings and the update banner probe /system/info on mount and treat
capability as advisory: the button always issues the request and shows the
server's actual refusal instead of sitting disabled after a failed probe.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Route admin CREATE/DROP DATABASE through a short-lived postgres.js maintenance
connection instead of spawning psql via execSync per call, and remove the
redundant DROP-before-CREATE (db names are pid+random, never pre-exist).
Cuts ~2 of 3 subprocess forks per test across ~55 tests; the slowest core
test file drops from ~90s under full-suite contention (32.6s->27s standalone)
with all 75 tests still green.
Fusion-Task-Id: FN-SLOW-TEST
Quick-add "Start" collapses create+promote into one request: it submits the
workflow id AND the post-intake `todo` column together, so the card lands in
`todo` having never sat in the workflow's manual intake column. The intake test
in task-creation.ts only matched `triage` or the resolved intake column, so the
card got generateSpecifiedPrompt — whose hard-coded boilerplate steps
("Implement the required changes") no planner ever wrote.
That stranded the card permanently: triage's todo-discovery admits a card only
when its PROMPT.md reads as a seed, so the placeholder spec was classified
"already planned" and never planned, while nothing could execute it either
(steps: []). It sat in Todo forever with no log line in any lane. Observed on
FN-8587.
Creates into `todo` on a manual-intake workflow (resolved intake is not the
legacy `triage`) now get the bootstrap seed. The pinned contract for a plain
direct create into todo on the default workflow — which intentionally keeps
generateSpecifiedPrompt — is untouched, and both create sites are fixed in step.
Also instrument the hold/release sweep, which had reasons but no timings:
per-task held duration reported on release, a per-sweep summary breaking out the
prefetch cost (a sequential await per non-archived task, so it scales with board
size rather than with held cards), and a warn when a sweep exceeds 2s — so a
"ready card doesn't move" delay can be attributed between poll cadence, sweep
cost, and a card genuinely queued on capacity.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Pressing Start on a Coding (Ideas) card only writes a column move — there is no
dispatch call in that path — so planning did not begin until the triage
processor's next timer tick, up to pollIntervalMs (15s default) later. The
"Started planning" toast was optimistic and the card just sat in Todo.
- Wake planning discovery on the store's task:updated/task:created event when a
task lands in todo/triage. Binding the wake to the store event rather than the
Start button covers every move surface (board drag, context menu, task detail,
List view, CLI, agent tools, POST /tasks/:id/move) by construction. The wake is
advisory: it only advances WHEN the poll runs, so every pause, seed-prompt,
dependency, and concurrency gate still applies.
- Admit a todo task whose PROMPT.md is missing instead of dropping it through a
silent `catch {}`. The scheduler KEEPS a candidate whose prompt it cannot read,
so such a card was invisible to planning while still visible to dispatch, with
no log line in either lane. Unreadable (non-ENOENT) prompts now log.
- Route the scheduler's dispatch filter through the shared isUnplannedSeedPrompt
predicate. Its open-coded strict bootstrap compare disagreed with triage on the
refinement-seed shape, leaving hold-release as the only thing between an
executor and a prompt containing just the operator's feedback text. The
predicate also normalizes line endings/trailing whitespace, so a CRLF or
trailing-newline round-trip no longer reclassifies an unplanned card as planned.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A legacy source column (todo/in-progress/...) validated moves only against the
closed VALID_TRANSITIONS map, which cannot know about a workflow-declared
column, so Todo -> Ideas was rejected even though the board drag pre-check and
context menu both offered it. Legacy sources now union VALID_TRANSITIONS with
the task's workflow-resolved adjacency, resolved lazily only when the legacy
table alone would reject. builtin:coding adjacency is unchanged.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
An operator can hold a raw Anthropic API key and a Claude subscription OAuth
login at once, and the raw key always won silently. A stale or revoked saved
key therefore shadowed a working subscription and failed every direct Anthropic
call with 401 invalid x-api-key, while both Settings cards still read Active.
Add the global anthropicAuthPreference setting ("api-key" default, preserving
the historical precedence, or "subscription"), read in resolveAnthropicRuntimeApiKey
straight from ~/.fusion/settings.json so it applies without a restart and needs
no settings plumbing through createFusionAuthStorage. Neither value removes a
source: with one credential configured, resolution reaches it either way.
Settings -> Authentication now names the credential in use on the two Anthropic
cards and renders the control, but only when both are actually connected.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The proactive status narration printed the internal 0-based step index, so the
final step of a 13-step task announced "Starting Step 12" next to a card
showing "12/13". Display now uses index + 1 in both the engine builders and
the store-side updateStep narration; the 0-based tool/PROMPT.md contract is
unchanged.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Binary Release (v0.73.0-beta.5 was fully red):
- bun compile: mark chromium-bidi external — playwright-core@1.60 (feature-video)
optionally requires it and bun fails closed on unresolvable requires.
- Windows desktop EXE: quote -c.publish.channel=beta in release.yml; PowerShell
tokenizes the bare flag into `-c` + a path and electron-builder ENOENTs on it.
Full suite (all 4 shards red from stale-test drift, no product bugs found):
- engine: align mock stores/assertions with atomic store.moveTaskIf dispatch
(#2371), the fail-closed non-empty PROMPT.md artifact gate (#2390), oldest-
first admission (FN-8453), alreadyClaimed graph routing (#2393), startStep
step projection (#2403/FN-8464), structured retry presentation (FN-8503),
provider-lane pause reasons (#2339), typed column-boundary entry (#2378),
Type.Integer in CAS document schemas (#2375), bounded model-registry refresh.
- engine-no-blocking-shellout: re-pin 17 drifted allowlist line numbers and drop
the stale REBASE_HEAD entry whose execSync was removed.
- core: schema-applier expectations track migrations 0033-0035 (96 tables) and
the synthetic 0000 fixture gains workflow_work_items/mission_contract_assertions;
work-item terminal state is "succeeded" post-#2378.
Known follow-up (not addressed here): self-healing starved-refinement escalation
bumps task.priority, which FN-8453 oldest-first admission no longer consults.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Reported bug (screenshot): deleting the task created from a plan left the
session permanently stuck on PLANNING_CREATED_TASK_MISSING — Retry create
replayed the same 409 forever. A linked task absent from the
include-archived scan (task-row authority; a successful scan proves
deletion, not a flaky read) now clears the stale linkage and creates a
fresh task, in both the create-task route and createTaskFromPlanSession;
a still-listed-but-unreadable task keeps failing closed.
Multi-agent review of fdd120232 (correctness/adversarial/reliability):
- P1: CLI planning sessions were memory-only — setAiSessionStore only ran
in the dashboard server, so --resume could never find a session across
invocations. New ensureDurablePlanningSessionStore wires the durable
AiSessionStore over the board store's public asyncLayer in runTaskPlan.
- P1: resume failures now THROW instead of process.exit (fn_task_plan
runs inside the pi host — an exit killed the whole agent session), and
a no-question resume requires an explicit refine focus (the provided
description) so merely resuming never rotates the epoch.
- P2: claim and finalize CAS gained the same expected-epoch WHERE guard
as reconcile, so a stale-epoch creator can no longer finalize an
old-epoch task onto a rotated session.
- Side-effect failures (documents, logEntry, validate, reconcile) are now
logged instead of swallowed; post-insert failures no longer mislabel
the just-created task alreadyCreated:true; the keep-refining readline
closes on thrown prompts and a failed refine after creation returns the
created task id with a resume hint; cross-process generating guard
added to createTaskFromPlanSession.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>