Release a settled retry token so a later planning error can start the next bounded attempt.
- Clear only the matching retry owner before scheduling its successor.
- Cover distinct stream errors after retry settlement.
- Add a patch changeset for the recovery fix.
Files changed:
.changeset/fn-8536-planning-retry.md | 7 +++++++
packages/dashboard/app/components/PlanningModeModal.tsx | 10 ++++++++++
.../__tests__/PlanningModeModal.planning-flow.test.tsx | 5 +++++
3 files changed, 22 insertions(+)
Fusion-Task-Id: FN-8536
Fusion-Task-Lineage: cccb0793-390e-4bdb-b459-939760e644bb
Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
After a hard host crash (SIGKILL, power loss), postmaster.pid survives with no
postmaster behind it. The optimistic join handed every subsequent boot a URL to
the dead port, so the dashboard could never start again without a manual pid
delete. Probe the recorded pid with signal 0: provably dead (ESRCH) rebuts the
live-lock presumption and the boot takes an owned start — PostgreSQL itself
re-validates and reclaims the stale lock file, so a recycled live pid keeps the
old join-then-fail behavior and a genuinely live postmaster still surfaces the
lock collision we already join on. EPERM counts as alive (fail-closed).
Verified end to end: real cluster started, postmaster SIGKILLed leaving the pid
file + interrupted WAL, fresh lifecycle detected the stale lock, ran an owned
start, and crash recovery preserved the marker row.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Reloading Planning re-entered the generation view ("Generating initial
plan…", Stop button, elapsed timer, 8s watchdog) while merely fetching a
persisted session. A new session_loading view state renders a neutral
"Loading session…" spinner during hydration; the generating view is
reserved for sessions the server reports as generating. Unrecognized
persisted session shapes now land in the retryable error view instead of
spinning forever.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The runner stage installed the application at /project, which was also the
documented bind-mount point — mounting a host project there shadowed the CLI
and the container exited with MODULE_NOT_FOUND on packages/cli/dist/bin.js.
- Install the app under /app and run the entrypoint by absolute path.
- Reserve /workspace (empty in the image, container workdir) as the project
mount point.
- Update docs/docker.md: mount at /workspace, and document that embedded
Postgres/global state lives in /home/node/.fusion with a named-volume
example so persistence actually captures it.
Verified: image builds; `docker run -v host:/workspace` boots, embedded
Postgres initializes, /api/health returns ok.
Fixes#2414
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
beta.4 follow-up from the issue thread: on an interrupted (not-cleanly-shut-down)
cluster, the elevated Windows launcher declared readiness on a bare TCP accept
while crash recovery still rejected every connection with 57P03, so
ensureDatabase failed and the start cleanup fast-shutdown the recovering
postmaster ~0.2s after launch; the retry then joined the instance it had just
told to stop, got ECONNREFUSED, and parked the dashboard in a dead shell. The
30s ".pgrunner sharing violation" stall was recovery's SyncDataDirectory fsync
walk hitting Fusion's own pgctl log inside the data dir.
- Move the pgctl runner dir to a sibling .pgrunner-<dataDirName> outside the
data dir (and sweep the legacy in-dataDir .pgrunner), so recovery's fsync
walk can never contend with the postmaster's inherited log handle.
- Ignore 57P03 recovery rejections in the elevated readiness fatal scan.
- Owned starts wait for the cluster to genuinely accept connections (retrying
57P03/socket errors, bounded by the start timeout) before ensureDatabase —
never stop a postmaster that is still in recovery.
- Join-path database verify retries the 57P03 recovery signal for up to 15s;
socket errors keep the instant optimistic-join contract for stale pids.
- startup-factory's joined-instance-unreachable retry backs off across ~15s
instead of a single 500ms attempt.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
## Summary
Wave 15 of package code organization.
### Peels
- `types/settings-scope.ts` — global/project settings (~2.2k lines)
- `types/archive-planning.ts` — archive, mesh/multi-project, planning
sessions
- `task-store/project-store-ops.ts` — rename of `remaining-ops-1` (last
numbered ops module)
### LOC
- `types.ts` ~5872 → ~3074
## Test plan
- [x] `@fusion/core` typecheck
- [ ] CI merge gate
**Stack:** this PR → #2397 → #2398
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Refactor**
* Reorganized and expanded the core public type surface into dedicated
modules for settings, archive/planning, board, tasks, todo lists, plugin
activation, and multi-project setup.
* Improved the browser-safe type exports to keep the public contracts
consistent.
* Updated internal project-level operation wiring to use the correct
project implementations.
* **Bug Fixes**
* Fixed a workflow creation test hook to inject the correct pre-insert
behavior for workflow-definition collision/allocator scenarios.
* **Chores**
* Refreshed internal headers and updated line-count baselines.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
## Summary
Restore full-suite green after latest `origin/main` landings. PR #2392
already merged; this branch starts from current main.
- **ExecutorStatusBar mobile utility test:** assert `Waiting` (product
rename from Queued) and that the In Review segment is gone.
- **AppearanceSection task-popup help:** assert FN-8478 fallback copy
(board deep-tab chips + List row/card + right-dock → movable popup).
- **Changeset format gate:** shorten
`mobile-board-pointercancel-settle.md` summary to ≤120 chars (was
blocking `pnpm test:gate` on main).
## Context
Latest main Full Suite run failed primarily on dashboard app quality
backfill:
- https://github.com/Runfusion/Fusion/actions/runs/29970703115
## Test plan
- [x] `pnpm test:gate` ×2
- [x] `utility-mobile.test.tsx` + `AppearanceSection.test.tsx`
- [x] `settings-default-descriptions.test.tsx`
- [ ] CI PR checks / Full Suite on this branch
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Improved mobile board column snapping when resting mid-screen,
flinging past columns, or making unintended swipes.
* Updated task status labels to show “Waiting,” “Running,” and
“Blocked.”
* Refined help text for opening tasks as popups, including click-target
and subsequent-task behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
## Summary
- Add `CHAT_CODEBASE_ACCURACY_GUIDANCE` so agent chat (direct +
multi-agent rooms) investigates the live checkout with tools before
answering code/architecture questions.
- Soften chat’s pure brevity default for repo questions: still lead
short, but keep real path/symbol evidence (parity with the investigation
pressure that makes Planning Mode more accurate).
- Cover the constant and assembly via unit tests; include a patch
changeset for `@runfusion/fusion`.
## Why
Users reported Planning Mode was more accurate about the codebase than
agent chat. Plan mode inherits the triage seam’s “read/grep first, name
real files” contract; chat only had a short helpful-assistant persona
plus a brevity default, so models often answered from priors.
## Test plan
- [x] `pnpm --filter @fusion/dashboard exec vitest run
src/__tests__/chat-system-prompt.test.ts
src/__tests__/chat-manager.test.ts`
- [ ] Manually ask agent chat a project-specific architecture question
and confirm it greps/reads before answering with real paths
- [ ] Confirm non-code chat still stays short/crisp
- [ ] Confirm multi-agent room responders also receive the new guidance
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Improvements**
* Improved repository/code answers by investigating the live codebase
first and prioritizing verified paths and symbols over speculation.
* Refined response-length behavior so code questions stay
evidence-focused, while non-code questions remain concise.
* Applied consistent accuracy guidance across both direct and room-based
conversations.
* **Bug Fixes**
* Prevented chat instructions from depending on unavailable mailbox
functionality, improving reliability for agentless/room flows.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
## Symptom
Users reported Planning Mode regularly failing with **"AI returned no
valid JSON.. Retry this planning session or start a new one."**, and
that leaving and returning to the interface mid-generation **duplicates
the generation — infinitely**. Every app-tab switch unmounts the
Planning view, so the leave/return path is the *normal* path, not an
edge case.
## Root causes
1. **Check-then-act turn admission.** `activeGenerations.has()` was
checked at turn entry, but the record was only created inside
`runGenerationWithTimeout`, after several awaits. Overlapping turn
entries (remounted view re-submitting, racing auto-retries, duplicate
start of an existing session) both passed the guard; the second
displaced the first, and the displaced teardown disposed the
**session-shared agent the surviving turn was actively prompting** —
which then read an empty assistant message and failed parse with "AI
returned no valid JSON".
2. **Per-mount auto-retry budget.** The client reset its 3-attempt
auto-retry budget on every mount, so each return to an errored session
re-ran a full-turn regeneration (agent rebuild + complete history
replay) — forever.
3. **SSE replay appended onto existing output.** Fresh stream
connections replay buffered thinking; the client pre-seeded from
persisted `thinkingOutput` (or kept prior output on silent reconnect)
and then appended the replay — visibly doubling the generation on every
reconnect. The 100-event buffer only held a suffix of a turn, forcing
that pre-seed.
4. **Raw `session.prompt()` at context limits.** A long interview that
overflowed the model's context window errored terminally, and auto-retry
replayed the full history into a fresh agent — overflowing again,
unrecoverably.
## Fix
- Synchronous per-session **turn reservation** shared by
`submitResponse`, `retrySession`, `startExistingSession`, and the
initial turn; losers get `GenerationInProgressError` instead of
displacing the winner. Duplicate starts of a generating session are
no-ops. Rewind aborts an active generation through its own teardown
first.
- Client auto-retry budget is **module-scoped per session** (survives
remounts); exhausted budget shows the error view instead of a stuck
spinner. Retry rejections for "already in progress" rejoin the live run.
- Fresh SSE connections **clear streamed output before the buffered
replay**; buffer deepened to a full turn (2000 events); rejoin paths
reconnect cleanly instead of seeding persisted thinking.
- All six planning prompt sites route through the engine's
**`promptWithFallback`**, recovering context-window overflows via
prompt/memory compaction and `session.compact()`.
- Cosmetic: no more doubled period in the retryable parse error message.
## Symptom Verification
- **Original symptom:** "AI returned no valid JSON" after answering
questions; generations duplicating on leave/return.
- **Exact reproduction:** concurrent turn entries on one session
(submit×2, retry×2, start-while-generating) — previously
displaced/disposed the live agent mid-prompt.
- **Assertion it is gone:** `planning-turn-admission.test.ts` asserts
exactly one turn is admitted per race, the winner completes with a
question and no session error, and the shared agent is never disposed;
`planning-context-compaction.test.ts` asserts every planning prompt
routes through `promptWithFallback` (signal forwarded) and that a
recovered context overflow leaves the turn healthy.
## Verification
- `vitest run` on all 7 planning server test files + the 2 new
regression files: **40/40 pass**.
- `routes-planning*` failures at main tip are pre-existing (identical
103/126 + 3/6 counts with and without this diff; main is mid-refactor on
route wiring). `planning-answered-question-reemit` 3 timeouts also
reproduce on clean main.
- `pnpm verify:fast` green; dashboard `tsconfig.json` +
`tsconfig.app.json` typechecks clean; eslint clean on touched files.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Prevented Planning Mode from duplicating generations and triggering
“AI returned no valid JSON” errors when leaving and re-entering mid-run.
* Made planning turn handling concurrency-safe and idempotent across
submit, retry, rewind, and duplicate start actions.
* Improved SSE reconnect recovery: clearer replay after reconnect, no
duplicated “thinking” output, and preserved auto-retry limits across
remounts.
* Improved long-context recovery via fallback prompting and cleaned up
retry error formatting.
* **Tests**
* Expanded coverage for concurrent Planning actions, reconnect replay,
context compaction, rewind behavior, and retry formatting; improved
parallel test-harness reliability.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
## Summary
Two-part fix for #2411 — Windows embedded PostgreSQL backends dying with
exception `0xC0000142` and taking the whole dashboard down.
### 1. Crash hardening + recovery (FN-8522)
- Child-only native `PATH` hardening so forked backends can always
resolve their runtime DLLs.
- Non-blocking `.pgrunner` log monitoring (shared read), eliminating the
self-inflicted ~30s `sharing violation` retry window at boot.
- Detection of the ordered 0xC0000142 shutdown sequence with a single
automatic restart of owned clusters on their resolved port, plus
operator diagnostics.
### 2. Platform-aware `max_connections` default (follow-up from
[operator
report](https://github.com/Runfusion/Fusion/issues/2411#issuecomment-5054900702))
On Windows every PostgreSQL connection is a separate process; the
embedded cluster's unconfigured `max_connections=500` cap lets backend
spawn bursts exhaust the non-interactive desktop heap, which kills
forked backends with exactly `0xC0000142`. The reporter confirmed
stability after lowering the cap.
- `embeddedPostgresMaxConnections` is now schema-unset so the server can
distinguish "operator never set it" from an explicit choice
(`getSettings()` merges schema defaults, which previously pinned 500
unconditionally and made the runtime fallback dead code).
- New `resolveEmbeddedMaxConnections()` resolves the unset default
platform-aware: **150 on win32, 500 elsewhere**. Explicit settings are
honored on every platform, clamped to [32, 2000] as before.
- Settings UI renders the cap empty ("auto") with platform-aware help
copy across all six locales.
- Fixed a latent reset bug this exposed: global "Reset this menu" wrote
`undefined` for undefined-default keys, which JSON serialization drops —
the stored value silently survived reset. Now uses null-as-delete.
## Testing
- New unit tests for `resolveEmbeddedMaxConnections` (platform defaults,
clamping, non-integer handling).
- Updated settings-defaults, default-descriptions, and SettingsModal
tests; embedded lifecycle + recovery coverage from FN-8522.
- `@fusion/core` builds clean; changesets included for both parts.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
## Summary
Gives dashboard integrators (plugin views, embedded panels, theming
tools) a supported way to match the dashboard's look and to layer
overlay UI correctly — instead of scraping computed styles and guessing
z-index values. This implements the CSS-token bridge slice of
`docs/proposals/2026-07-01-dashboard-theme-plugin-system.md`.
Two additions, both inert unless used:
1. **Documented theme-token contract.** A "Theme tokens" section in
`docs/dashboard-guide.md` (referenced from `docs/PLUGIN_AUTHORING.md`)
declares the stable set of CSS custom properties — colors, surfaces,
status colors — that integrators may read. Tokens resolve to raw color
strings (e.g. `#161b22`) in every theme, including the newer ones. A
sync test (`theme-token-contract-docs.test.ts`) parses the doc's token
table and asserts each documented token has a real definition in the
dashboard CSS, so the contract cannot silently drift from the code.
2. **Overlay layering surface.** Overlay-style UI (palettes, pickers,
floating panels) currently has no supported way to sit above the
floating-window stack — the effective max z-index is runtime state
inside `floatingWindowStack.ts`. This PR exposes it:
- `--fusion-max-z` on `:root` — kept in sync by `floatingWindowStack`
(written at module load and after every `nextFloatingZ()` claim), so it
always reflects the true top of the dashboard-managed stack. Boot/floor
value is `11001`, chosen to clear the highest statically-declared layer
(the body-portaled model-combobox dropdown at `z-index: 11000`).
- `#plugin-overlay-root` — an empty, `pointer-events: none` sibling of
`#root` stacked at `calc(var(--fusion-max-z) + 1)`. React never renders
into it, so it is hydration-safe; integrators portal into it and
re-enable pointer events on their own elements.
- The layer bands (base UI / floating windows / toasts / dropdown /
overlay root) are documented in `styles.css` and the guide, and a guard
test (`dashboard-max-z-guard.test.ts`) scans the structural + component
CSS and fails if any static `z-index` is ever introduced above the floor
— keeping the contract honest as the codebase evolves.
## Behavior
No visual or behavioral change for existing users: `floatingWindowStack`
still returns the same values from `nextFloatingZ()`; the overlay root
is empty and click-through; tokens were already defined — this only
documents and guards them.
## Tests
- `theme-token-contract-docs.test.ts` — docs ↔ CSS sync
(non-tautological: anchored matching against real definitions).
- `floatingWindowStack.max-z.test.ts` — `--fusion-max-z` boot value and
live tracking as the stack claims z-indexes.
- `dashboard-max-z-guard.test.ts` — no static dashboard z-index above
the floor (decorative `public/theme-data.css` INT_MAX scanline overlay
deliberately excluded; it's non-interactive grain, documented in the
test).
- Changeset included (`minor`, `category: feature`). Typecheck clean.
## Open question for maintainers
The token is named `--fusion-max-z`. The existing scale uses `--z-*`
names (`--z-dropdown`, `--z-modal`) on a lower band — happy to rename to
`--z-max` / `--z-plugin-overlay` or anything that fits your convention;
the name is the only bikeshed here, the sync mechanism is independent of
it.
## AI assistance disclosure
Parts of this change were authored with AI assistance (Anthropic's
Claude); the commit carries a `Co-authored-by` trailer accordingly.
Everything was human-reviewed before submission, and the test suite and
typecheck were run locally against the current `main`.
If squash-merging with a rewritten message, please keep the attribution:
```
Co-authored-by: Claude <noreply@anthropic.com>
```
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **New Features**
- Added stable dashboard theme tokens for consistent plugin/integration
styling.
- Introduced a dedicated plugin overlay mount point with click-through
defaults and an overlay stacking ceiling.
- Overlay z-index now stays in sync with floating window layering
automatically.
- **Documentation**
- Added an explicit stable “theme token contract” and “overlay layering
contract,” including interaction and z-index usage rules and deprecation
expectations.
- **Bug Fixes**
- Improved reliability of plugin overlay stacking so overlay content
renders above intended dashboard layers.
- **Tests**
- Added guards validating CSS z-index ceilings and enforcing the
documented theme token contract.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Co-authored-by: Claude <noreply@anthropic.com>
The jump-to-latest button is centered via transform: translateX(-50%), and
the global .btn:active scale transform replaced it wholesale on mousedown,
shifting the button out from under the cursor mid-click. Compose both
transforms on :active (same fix as the DevServer new-logs button).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Plugin-defined HTTP routes were a boot-time snapshot of the launch
project's PluginLoader, while dashboard views/UI slots resolve a
project-scoped loader live per request. Two failure modes survived the
961edf214 no-engine mount fix: a plugin enabled after boot rendered its
view while every API route 404'd until restart, and a plugin enabled
only in a non-launch project never got routes mounted at all (Compound
Engineering "Failed to load sessions: Not found" on v0.73.0-beta.3).
Routes are now dispatched per request through the same
getProjectPluginLoader cache (moved from the plugins registrar into
routes/context.ts) that serves dashboard-views and enable/disable, with
the host loader + PluginRunner tables unioned in (project entries win,
loader beats runner). The compiled dispatch sub-router is cached per
resolved loader and rebuilt only when the route signature changes, so
views and routes agree by construction.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Code-review findings on 085f7b99c: a bare date like 2026-07-22 qualified
as a distinctive slug, so unrelated failure reports quoting the same date
silently converged (and the date outranked a real file-path anchor in the
sorted-first pick); reject slugs whose segments are all hex/numeric.
Also render non-http(s) sourceMetadata.issueUrl values as plain text to
block javascript:-scheme links from API-supplied metadata.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Widen computeCrossParentDiagnosticClaim so repair tasks phrased as
'exceeds limit / oversized / blocking X / so X passes' converge on a
file-path or distinctive-slug anchor at creation time (FN-8510/8511/
8513/8514 incident: four executors on unrelated parents filed the same
oversized-changeset follow-up and none deduped before triage).
Add a Provenance section to the Task Detail Stats tab showing source
type, parent task, creating agent, imported-issue link, and the triage
near-duplicate marker.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The KTD-8 adoption sweep runs on every store open, and a DB with active tasks
never records the drained marker, so it re-runs constantly. Its resume-graph
mapping for 'planning' cleared FN-8504's freshly written live planner status
~100ms after triage claimed it (audit: task:reconcile-legacy-adoption,
priorStatus 'planning'), leaving a live replan planner rendered as an idle
READY card and invisible to every Running count.
Generalize the FN-8498 needs-replan fix: any status with a live post-cutover
writer is preserved — planning (triage's stale-planning sweep owns crash
recovery), queued (scheduler re-evaluates each poll), merging/merging-pr/
merging-fix (self-healing stale-merge recovery), stuck-killed (restart-
recovery coordinator). Only writer-less statuses (plan-review-unavailable,
triaged) keep resume-graph.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A Todo card parked in the durable needs-replan stage keeps its REVISING badge
and activity chrome (FN-8494), but the header only counted live agents, so it
read 0/2 under a glowing card. Union the shared Running predicate with the
card's own activity predicate (same globalPaused/stuck gates) so the header
equals the number of visibly active cards. Footer Running and admission keep
live-agent-only truth: a parked replan must not consume concurrency capacity.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Align the mobile task-detail Move action with the footer edge.
- Let the footer spacer absorb surplus mobile width instead of the Move dropdown.
- Cover canonical footer order and triage actions-absent behavior.
- Add a patch changeset for the mobile footer fix.
Files changed:
.changeset/fn-8501-mobile-footer-alignment.md | 7 ++++
.../dashboard/app/components/TaskDetailModal.css | 8 ++++-
...etailModal.responsive-and-dependencies.test.tsx | 40 ++++++++++++++++++++++
3 files changed, 54 insertions(+), 1 deletion(-)
Fusion-Task-Id: FN-8501
Fusion-Task-Lineage: a87bb5af-d16e-4ca8-9111-3168c8440f14
Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
Two mobile board gesture fixes: (1) finger travel only counts as horizontal
pan intent when it dominates the vertical axis, so scrolling a column's card
list with incidental diagonal drift no longer swipes to another column; (2)
the commit-one-column paging clamp applies only to gestures begun at rest
centered on a column — a tap-to-stop mid-momentum followed by a drag settles
on the nearest column at the drag's landing point instead of being forced a
column past it by the interrupted scroll's origin.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
An In Review column with one MERGING task and one live CODE REVIEW task showed
1/2 processing: gate sessions run with task.status left null, so the shared
isRunningAgentTask predicate only saw the merge-pipeline statuses. Count a
pending workflow-step-result lease (the durable live-gate signal; FN-8492
fails orphaned ones) as Running on any unpaused, non-terminal row. Covers
plan review, code review, browser verification, post-merge verification, and
custom optional steps across column headers, footer stats, admission, and CLI.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The directional settle always paged one past the nearest column once the
viewport center crossed its center, so a fling that decelerated with a column
mostly on screen was pushed a full extra column. Settle now lands on the
nearest column, clamped to at least one column of progress from the gesture's
origin column — short swipes still commit forward and the settle never moves
against travel.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
iOS/Android fire pointercancel when the native pan claims a touch while the
touch stream keeps flowing. Treating it as gesture end either orphaned the
gesture (drag released mid-screen rested between columns until the next tap)
or armed the idle settle with the finger still down (slow scrolls at the edge
columns glitched and snapped back). Track the live touch sequence and ignore
pointercancel while it is active; touchend stays the real finger lift.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Code-review follow-up on 4413699de. Deleting an orphaned pending review
entry was a severity inversion: the merge gate blocks on pending/failed
results, not on an enabled step with NO result, so deletion silently
satisfied the gate and the task merged with its review skipped (verified
live: FN-8492 landed on main without Code Review re-running). Orphans are
now rewritten to status:"failed" — the gate stays closed and the
failed-pre-merge-steps recovery / FN-7720 operator-bypass paths own the
re-run decision.
Also from review: the sweep now runs in periodic maintenance too (a step
session can die without a restart), skips executor-owned in-progress rows
(resume is deferred ~30s at startup, so their liveness is unprovable when
startup recovery runs), re-reads the row immediately before the write so
the whole-array update cannot clobber a fresh lease, counts recovery on
the successful mutation rather than after the audit emit, and the new
audit event literal is registered in DatabaseMutationType (cast dropped).
Tests now cover all three liveness-triple legs, >500-row pagination,
in-progress skip, per-task write-failure isolation, and the never-delete
invariant; the needs-replan adoption row moved under a preserve-group
header.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>