Extends the doc merged in #3255 with two more instances of the same pattern, both found this session, **neither involving a ratchet**. Four instances now, from four unrelated directions: | what was read as "pass" | what the green actually meant | | --- | --- | | `node scripts/check-*.mjs` exits 0 | report-only mode — the failure path needs `--strict` | | a census reports 0 for a new file | the file is untracked, so it was never scanned | | a backgrounded `cmd > log; grep …` reports exit 0 | that is `grep`'s status; the suite inside had 8 failures | | a rebased branch's tests pass | the rebase never started, so it ran on the **old** base | The two new ones are worth writing down because they are not about tooling anyone built here — they are about how results are read. **Exit codes belong to the last command in the pipeline.** A backgrounded `run_tests > log 2>&1; echo done; grep X log` exits with `grep`'s status, so the harness reported "completed, exit code 0" for a dashboard suite that had 8 failures. I nearly recorded that suite as green. Read the summary out of the log; never infer a suite's result from a wrapper's exit code. **A failed rebase leaves you on the old base, and the tests still pass there.** `git rebase` refused with `cannot rebase: You have unstaged changes`, so the branch never moved. `git diff origin/main` then listed 20+ files including other workers' commits — which reads exactly like my branch had reverted their work — and a full test run on that tree came back green. Both signals were true about a tree nobody cared about. ``` git merge-base --is-ancestor origin/main HEAD ``` said STALE while the tests said pass. That is the only check that separates the two, and it belongs before any claim of "verified on current main". The shared tell, stated once: **a result too clean, or too alarming, for what changed.** Every probe shape passing including ones that obviously should not; a two-file branch appearing to revert twenty. When the answer does not fit the size of the question, find out what was actually measured before believing it. ## Verification Docs only; no code paths change. `lifecycle-columns`, `move-target-literals`, `inert-sync-lanes`, `quarantine-ledger` all exit 0. No changeset — AGENTS.md excludes internal docs. **Pre-existing red, not from this branch:** `check:fnxc-future-dates` currently fails on main from a `2026-08-01-00:50` stamp in `packages/core/src/task-store/lifecycle-ops.ts` (commit `e52da740a5`) — a timezone-ahead clock writing tomorrow's date, at 23:45 UTC. Already claimed by **#3269 and #3270**, so I have not touched it; flagging only so this branch's CI result is not misattributed. It is the same recurring class this doc's sibling rule addresses: take the stamp from `date -u`, not the local clock. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Added guidance for identifying misleadingly successful CI and test results. * Documented checks for report-only runs, untracked files, masked failures, and tests running on an outdated code base. * Included recommendations for reviewing logs and verifying branch ancestry. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Fusion Documentation
Fusion is an AI-orchestrated task board that turns ideas into reviewed, merged code using a structured workflow: planning → todo → in-progress → in-review → done.
Quick Start
Start the local dashboard with pnpm dev dashboard, then create your first task from the board or CLI.
For a full walkthrough (installation, onboarding, first task, and daily workflow basics):
Documentation Index
Getting Started
| Guide | Description |
|---|---|
| Getting Started | Installation, first-run, first task, and daily workflow basics |
| Dashboard Guide | Board/list views, left/right sidebar navigation, Artifacts, Import Tasks, chat, workflow selection/editor, terminal, git manager, files, planning, and UI tools |
| CLI Reference | Complete fn command reference with subcommands, flags, and examples |
| Remote Access | Operator runbook for Tailscale/Cloudflare setup, tokenized login links, security caveats, and troubleshooting |
| Native Shell Connection Guide | Canonical mobile/desktop shell onboarding, profile management, QR/manual setup, and remote handoff behavior |
Task & Project Management
| Guide | Description |
|---|---|
| Task Management | Task creation modes, lifecycle, prompt specs, comments, archiving, and GitHub integration |
| Todo View | Canonical guide for the experimental Todo View, including enablement, usage, API routes, and storage |
| Missions | Mission hierarchy, planning flow, activation, progress tracking, and autopilot behavior |
| Goals Refinement Gate | Evidence gate for activating the conditional post-v1 goals refinement slice only after real usage pain is documented |
| Goals Refinement Evidence Pack | Structured observation template and two-observation threshold for conditional Slice 4 activation requests |
| Research | Research runs, provider setup, dashboard/CLI usage, findings, exports, and task integration |
| Research View UX Spec | Canonical layout and capability-state messaging spec for the Research dashboard view (FN-4138, informs FN-4134/FN-4135) |
| Workflow Steps | Workflow overview, built-in workflow catalog, per-task selection, runtime semantics, reusable quality gates, templates, phases, and execution results |
| Workflow Editor | Visual workflow editor guide for opening, viewing, authoring, validating, importing/exporting, custom fields/columns/settings, and tuning workflows |
| Custom Workflow Reliability Acceptance Map | End-to-end reliability acceptance criteria for custom workflow authoring, selection, execution, recovery, restart durability, and deferred journeys |
| Custom Non-Coding Workflows MVP Spec | MVP framing for user-authored non-coding workflows, lifecycle mapping, metrics, and risk checklist |
| Task Evaluations | Eval scoring contract, evidence persistence, score categories, and evaluation pipeline |
| Multi-Project | Central registry architecture, project management, isolation modes, and migration paths |
Configuration & Agents
| Settings Reference | Global/project settings, workflow setting values, model/fallback lane hierarchy, defaults, and API endpoints |
| MCP | Model Context Protocol server configuration, secret references, validation, CLI, dashboard, and import/export workflows |
| Agents | Agent management, presets, prompts, heartbeat behavior, spawning, and mailbox workflows |
| Planner Oversight (see Settings Reference, Dashboard Guide, Architecture) | Workflow-native oversight levels (off/observe/steer/autonomous), per-task overrides, notification verbosity, the human-confirmation gate on merge/PR and destructive actions, and the Task Detail overseer controls/Intervention Timeline |
Architecture & Development
| Guide | Description |
|---|---|
| Architecture | System architecture, package layout, storage model, and engine execution flow |
Secrets Store (SecretsStore) |
Core encrypted secret subsystem overview: scopes, AES-256-GCM at-rest model, policy semantics, and public store API surface |
| Dashboard Real-Time | Canonical event-stream architecture contract (shared /api/events bus + dedicated stream boundaries), with project/node scoping, reconnect/cleanup behavior, and realtime pitfalls |
| Storage | PostgreSQL runtime storage, archive, migration compatibility, and file-backed payloads |
| DAG Architecture Deliverables | Milestone A DAG architecture documents plus Milestone B prototype scaffold docs (schema migration plan, DagCoordinator design, implementation checklist) |
| Dev Server Module Audit | Analysis of parallel dashboard dev-server module families, production wiring, and consolidation guidance |
| Shared Cluster Protocol | Shared PostgreSQL multi-node contract: claims/leases, membership, auth, and retired multi-leader mesh replication |
| Signals Connectors | HMAC-signed external signal connectors for setup, payload mapping, and security notes across Sentry, Datadog, PagerDuty, and generic webhooks |
| Multi-Project Sequencing and Dependency Analysis | Sequencing guidance for FN-3448/FN-3449/FN-3503/FN-3182, including identity boundaries and recommended board dependency edges |
| Contributing | Local development setup, testing, release flow, and contributor conventions |
| Docker | Container builds, deployment, and persistence configuration |
| Code Signing | macOS and Windows code signing configuration for release binaries |
| Diagnostics | Engine diagnostic logging subsystems, structured log keys, and key diagnostic points catalog |
| Sandbox Backends | Pluggable sandbox backends for executor command isolation (bubblewrap, spawn-based) |
| Secrets | Encrypted secrets storage, per-secret access policies, scopes, and agent tool wiring |
| Testing | Full testing lanes, worker fanout guidance, test taxonomy, and file organization |
| Real iOS Safari Acceptance Surface | Provisioning runbook and harness usage for terminal verification gates on physical or cloud real-iOS Safari |
| Solutions Catalog | Documented solutions to past problems (bugs, architecture patterns, best practices) organized by category |
| Localization Contributing Guide | Conventions for contributing translations, locale file structure, and i18n tooling |
| Mobile | Capacitor/PWA mobile development setup and workflow |
Plugins
| Guide | Description |
|---|---|
| Plugin Management | End-user guide for discovering, installing, enabling, configuring, updating, uninstalling, and troubleshooting Fusion plugins |
| Plugin Authoring | Developer guide for building Fusion plugins (manifest, SDK hooks, routes, UI/runtime contributions) |
| Even Realities Glasses Plugin | Task-focused Even Realities glasses bridge with quick capture, polling notifications, and agent actions |
| Reports Plugin | Reports plugin rendering, export, standalone HTML generation, and section configuration |
| Even Realities Plugin API | Even Realities plugin API endpoint reference and test coverage matrix |
| Memory Plugin Contract | Pluggable memory backend architecture, interface contract, and migration strategy |
| Compound Engineering Plugin | CE workflow dashboard surface: artifact hub, interactive sessions, work→board bridge, and bidirectional sync |
| External Plugin Authoring | Step-by-step guide for authoring plugins using an installed fn CLI (no monorepo access needed) |
| External Plugin Proof-Point Runbook | Repeatable release-validation runbook for proving an external plugin runs against a published Fusion CLI build |
Audit Reports
| Report | Description |
|---|---|
| Test Feedback-Loop Baseline | Weekly FN-6612 signal-per-second baseline for gate/test wall-time, slowest files, and quarantine trends |
| Test Value Audit | Heuristic test-value audit generated by scripts/test-value-audit.mjs to support human deletion and review decisions |
| Test Velocity Baseline | Weekly feedback-loop velocity baseline for merge-gate, boot-smoke, changed-test, and quarantine metrics |
| UX Audit Report | Comprehensive UX audit with prioritized recommendations for dashboard improvements |
| Codebase Improvement Audit | Evidence-based technical debt and reliability gap audit with prioritized recommendations |
| Gap Analysis | System completeness analysis comparing Fusion to Paperclip feature set |
| Permanent Agent Heartbeat Playbooks | Worked manager/IC/message/blocked/no-task heartbeat scenarios and anti-patterns |
| Agent Sandbox Research | Research on agent isolation, capability enforcement, and sandboxing approaches |
| Even Realities Integration Research (FN-3737) | Research summary and recommended integration topology for Even Realities glasses + Fusion |
| pi-autoresearch Analysis for Fusion Port | Upstream architecture/license analysis and Fusion integration mapping for autoresearch capabilities |
| pi-autoresearch Audit vs Fusion Research | Audit comparing Fusion's research subsystem against upstream pi-autoresearch capabilities and parity gaps (FN-4136) |
| Research Hardening Preflight Baseline | Verified research subsystem baseline, lifecycle contracts, and hardening pressure points |
| Test Audit Report | Test coverage and effectiveness audit with recommendations |
| Skipped Test Inventory | Current intentional test-skip inventory and reconciliation status for older skip follow-ups |
| Dev Server Module Boundary Audit | Boundary/ownership audit for parallel dev-server-* vs devserver-* dashboard modules and FN-2212 prioritization guidance |
| spawn_agent Approval Evaluation (FN-3973) | Decision to keep fn_spawn_agent under generic action-gate governance rather than durable agent provisioning policy |
| Task Lineage Reconciliation Notes | Historical task-ID reuse patterns, confidence semantics for commit attribution, and reconciliation methodology (FN-3953, FN-3998) |
| Dashboard Load Performance (historical) | Pre-cutover SQLite index analysis retained for performance archaeology |
| CLI Printing Press Plugin Design | Architecture design for the CLI printing press bundled plugin (FN-3762) |
| CLI Printing Press Research | Upstream cli-printing-press analysis and Fusion integration mapping (FN-3761) |
| Research vs Experiment Session Naming Decision | Naming decision record: hybrid approach retaining research_* for cited-search/synthesis and adding experiment_session_* for upstream parity (FN-4223) |
| Experiment Executor Design | Experiment executor architecture: lifecycle, run state machine, and worktree isolation model |
| Experiment Finalize Flow | Experiment finalize contract: branch grouping, dry-run planning, and session completion semantics |
| Experiment Session Model | Experiment session data model: state transitions, iteration tracking, and persisted run state |
| Experiment Session MVP Spec | MVP specification for the experiment session feature: scope, invariants, and delivery milestones |
| Sandbox Options Research (FN-4635) | Pluggable sandbox options research: threat model, backend evaluation, and spawn-based isolation design |
| Triage Duplicate Detection Postmortem | Postmortem on duplicate task detection gaps and scheduler dedup hardening |
| Multi-Node Runtime Readiness (FN-4814) | Runtime readiness assessment for multi-node distributed coordination |
| Distributed Multi-Node Coordination Gap (FN-4819) | Gap analysis for distributed multi-node agent coordination and cross-node task assignment |
| Cross-Node Assignment Wake Contract (FN-4824) | Contract specification for cross-node task assignment wake signaling |
| Multi-Node Coordination Validation Findings (FN-4820) | Validation findings from multi-node coordination testing and edge-case analysis |
| Secrets Sync Auth Parity Review (FN-4886) | Review of node secrets sync API authentication parity and security boundaries |
| Test Speed Audit (FN-5048) | Measured baseline test performance, offender list, and optimization priorities |
| Soft-Delete Verification Matrix | Authoritative checklist for the FN-5105 → FN-5143 soft-delete stream: scenario × layer coverage |
| Self-Healing Backward Move Audit | Audit of self-healing backward-move safety checks and edge-case validation |
| Workflow Policy Ownership Map | U1 characterization map classifying production merge, retry, scheduling, and recovery policy branches before workflow-policy migration cutover |
| Test-Speed Baseline (2026-06-03) | Measured per-file test timing baseline and optimization targets (successor to FN-5048 audit) |
| ACP Runtime Contract | Agent Client Protocol plugin launch/readiness contract and failure taxonomy |
| ACP MCP Passthrough & Permission Forwarding Upstream Sponsorship (FN-6475) | Ready-to-file upstream sponsorship for claude-code-cli-acp ACP session/new.mcpServers passthrough and permission-gate traversal; Route A remains NOT GO until proven |
| Mission Completion Gate Contract | Decision record for mission completion gate invariants and acceptance flow |
| Lost-Work Tasks Incident (2026-05-23) | Incident catalog of 9 lost-work tasks from no-op finalize and reuse-handoff bugs | | GitLab Parity Inventory (FN-7421) | Implementation map for first-class GitLab support: import, linked issue tracking, comments, auth/settings UI, CLI/extension, and Command Center surfaces to mirror or explicitly exclude | | PostgreSQL Runtime Cutover Review (2026-07-14) | Current end-to-end authority inventory, intentional legacy SQLite readers, deployment contract, and verification record | | SQLite → PostgreSQL Migration Review (2026-06-26, historical) | Historical multi-agent review of the incomplete migration branch and its original findings | | Dashboard Theme & UI Plugin System Proposal (2026-07-01) | Feasibility-spike proposal for a controlled dashboard theme/UI shell extension point sharing one backend source of truth | | Full-loop Agent Tool-Surface Audit and Delivery Plan | Source-grounded audit of engine-agent and dashboard chat tool factories, gap analysis for mission hierarchy integration, and delivery plan (FN-8280) | | Dashboard Modal Inventory | Canonical classification of all 45 dashboard modal surfaces (classes A–D) with file:line evidence, FloatingWindow migration targets, and the shared migration contract (FN-8605 → FN-8617) | | Workflow-Owned Lifecycle Closing Verification | Closing-bar verification runbook and recorded pass history for the workflow-owned lifecycle cutover programme — gate, verify:fast, E2E families, and census |
External Resources
- GitHub repository: https://github.com/Runfusion/Fusion
- npm package: https://www.npmjs.com/package/@runfusion/fusion
- pi agent framework: https://github.com/earendil-works/pi
Suggested Reading Paths
- New user: Getting Started → Dashboard Guide → Task Management
- Workflow author: Dashboard Guide → Workflow Editor → Workflow Steps → Settings Reference
- Power user / automation owner: Settings Reference → Workflow Steps → Agents → Planner Oversight (Settings Reference § Workflow Settings)
- Maintainer / contributor: Architecture → Multi-Project → Contributing
