**The ratchet was advisory.** `scripts/lifecycle-column-census.mjs`
existed only as `pnpm census:lifecycle-columns` — without `--strict` —
and **no workflow invoked it**. Nothing has ever compared the tree to
the baseline. Every "the baseline ratchet holds them" assumption in this
program rested on a check that does not run.
That explains both classes of hole:
**1. Three PRs lowered counts without re-recording,** leaving allowances
the deleted guards could return through while every check stayed green.
I've tightened them across #2593 and earlier PRs, but nothing stops the
next one.
**2. #2621 GREW the count while its own title claimed "count 0 → 0".**
It added `column === "triage"` and `column === "todo"` at
`register-task-workflow-routes.ts:2681`, taking that file to **23
against an allowance of 22**. It landed unchallenged. This is the
failure mode the ratchet exists to prevent, and it happened *inside this
program*, in a PR that asserted the opposite.
## The change
Adds `check:lifecycle-columns` (the census with `--strict`) to the
`pr-checks.yml` lint job, next to `check:changesets` and
`check:routes-modular` — the established pattern. **~1.8s over ~1950
files**, so this is not a slow-test addition.
## Proven to fail, in both directions
A guard that reports success without checking anything is worse than no
guard, so:
| injected defect | result |
|---|---|
| `const __probe = (c: string) => c === "triage"` added to `moves.ts` |
`count ROSE — moves.ts: 39 -> 40`, exit 1 |
| run against main's current baseline | exit 1 on
`mission-feature-sync.ts: allows 5, tree has 0` |
Both reverted; exit 0 restored. Note the second row: **this check is RED
on main right now**, which is the point.
## Merge order
**Stacked on #2593**, which carries the `DELIBERATE-LITERAL` marker for
the #2621 site (a v1 IR declares no roles, so no trait can answer that
question) plus the baseline re-record. Standalone on main this PR is red
— correctly. **Merge #2593 first**, then this.
I stacked rather than duplicating those two edits because I already
caused one conflict today by appending related content from two
branches, and #2651 merged a correction ahead of the section it
corrected. Same-content edits in two PRs is the same mistake.
## Census
Unchanged by this PR: **776 total, triage 5, reviewed 16** — it adds no
guards and converts none. It only makes the numbers enforceable.
## For the fleet
This should land before the 776-guard fleet launches. The brief says
"the baseline ratchet must shrink by exactly the converted count" —
until now nothing verified that claim, so a batch worker could report a
shrink that did not happen, or grow the count while converting, and CI
would agree.
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Completes `proving-a-code-path-actually-runs.md` (merged as #2642) with
the rule its own author broke three times while writing it. **Docs
only.**
## Why this belongs in that document rather than a new one
Findings 1-5 are about proving **your own** claim: does this path run,
can this test fail, is this negative result observable. Finding 6 is the
mirror image — the claims we make against **other people's** work — and
it is the same underlying error pointed outward. Splitting them would
let a reader take the first five as "be rigorous about my code" and miss
that the identical discipline applies when reviewing someone else's.
## The three cases, all mine, all in one day
| What I claimed | What was actually true |
|---|---|
| The census undercounts triage guards, 13 vs 10 | `summarize()` counts
`byColumnId` only for `kind === "column"`. My patched counter summed
`role`, `status` and `deliberate` too. The three "missing" ones were
exactly the ones it classifies correctly — and I reported this against
the instrument the program had just adopted as authoritative. |
| `resolvePlannerLanesForTask` silently disables two recovery paths for
legacy cards — escalated across four messages | The file's own header
had already reasoned it through and documented why that answer is
correct. And `TaskStore` implements `getTaskWorkflowSelectionAsync`,
which the resolver prefers — so real projects never take the path my `{
getTask }`-only probe forced. |
| `executor.ts` is clean of triage guards | A receiver-specific grep
missed three under `from` and `originColumn`. Same error one step
earlier: trusting a reconstruction of the thing instead of the thing. |
Every one was: reconstruct behaviour from outside → compare to actual
output → find a difference → report a defect, **without reading the
implementation.**
## The rules it adds
- Read the implementation and its header comment before reporting
anything as wrong. On this codebase the reasoning is usually already
written down, and the FNXC note frequently answers the exact objection —
twice today it answered mine verbatim.
- **A fixture is not a measurement of production.** When a probe and the
real system disagree, suspect the probe: ask what it had to stub, and
whether production ever supplies that shape.
- Retract precisely and immediately. A false defect report against
shared infrastructure costs more than the bug would have — it sends
people to verify something already correct, and spends the credibility
needed for the next report that is real.
Also updates the count in the intro (five → six) and adds an
`applies_when` entry so the doc surfaces for "about to report a tool as
defective", which is when it is needed and not when someone is already
debugging.
`pnpm lint` clean. No changeset — internal documentation.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two findings, no behavior change. Both are about **recorded reasoning
that was wrong** — the kind that sends the next person the wrong way.
## 1. The coding-ideas column collapse does not work (IR change
reverted)
I implemented it — deleted `ideas`, moved its `intake`/`autoTriage:
false` onto Planning, repointed the `start` anchor, updated the IR
suites to the merged shape (they went green, 44/44). Then the wider
suites failed and showed why it cannot work.
**The manual gate IS the column boundary.** `replan-target.ts` names the
discriminator in its own comment: *"The real discriminator is which lane
the triage service SCANS, which depends on the intake column's
`autoTriage` config."* So `ideas` is unscanned, `todo` is scanned, and
"promote" means moving the card from one into the other. Merge them and
one column must be both:
| if… | consequence |
|---|---|
| `autoTriage: false` wins | never scanned → nothing is ever planned →
the capacity hold releases an **unplanned** card into `in-progress`,
violating FN-7648 |
| scanning wins | `autoTriage: false` is meaningless → the manual gate
is gone → the preset duplicates the default Coding workflow |
**8 tests fail, and they are not fixtures** — they encode the promotion
flow itself, e.g. `store-create-intake-column.test.ts` › *"promotes an
Ideas-parked task to todo without planning it (still bootstrap-stub
PROMPT.md)"*. Rewriting them would have meant inventing what "promote"
means with no destination column, which is how a broken flow gets
blessed by a green suite.
**What it would actually take:** a promoted flag the triage scan reads,
so one column can hold both "not yet promoted" and "being planned". That
is a new lifecycle signal, not a column merge — the same shape as the
deferred `needs-replan` follow-up. Happy to scope it.
**I also corrected my own earlier checklist** in this doc, which said to
delete the now-dead `isUnplannedStartCreate` arm. Wrong: `autoTriage` is
a general trait field (`builtin-traits.ts`), so any custom workflow can
declare a manual intake with `intake !== hold`. The arm is dead only for
this preset.
## 2. `replan-target.ts` recorded the U11 merge backwards
The note claimed U11 deletes `todo` and keeps `triage`. It is the
reverse — Shape B kept the id `todo` and deleted `triage`, precisely so
the ~120 `column === "todo"` guards kept their meaning and no data
migration shipped. The default lineage now declares `todo, in-progress,
in-review, done, archived`.
The lookups are correct today, but **for the opposite reason to the one
recorded**: the default lineage falls *through* the `triage` lookup and
lands on `todo`, its merged planning column. `triage` still matches the
workflows that genuinely declare it (Lead generation, PR review).
Also flagged without changing (it would be a behavior change): the
`return "triage"` fallbacks on the no-match and throw paths name a
column the default lineage no longer declares, so a workflow with
neither `triage` nor `todo` gets a nonexistent target.
## Census
**Unchanged: 781 total, triage 5.** This PR adds no guards and converts
none — `workflowHasColumn(ir, "triage")` is a call argument, not a
comparison, so it is outside what the census counts either way.
## Verification
41/41 engine replan-target suites (including the existing
`replan-target-merged-planning-column` suite that covers the corrected
behavior) · engine typecheck clean · the reverted IR restores the tree
to main's content for those three files, verified by `git checkout --`.
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Durable write-up of U8's verification findings. **Docs only — no code
change, no CI risk beyond lint.**
These currently exist only in PR bodies, which nobody greps.
`docs/solutions/` is where this project keeps exactly this kind of
thing, and every one of the five will recur: the handler-pair shape and
the resolved-vs-guessed fork both have more call sites than U8 touched.
## The five
1. **Two prompt-node handlers exist; only one runs.**
`createDefaultNodeHandlers` prefers the primitives handler whenever
`deps.primitives` is set, and `executeWorkflowGraph` always sets it — so
every seam entry in `createAuthoritativeWorkflowSeams` is unreachable
for prompt nodes. A lifecycle announcement sat there through two PRs. It
type-checked and its unit tests passed, because a seam-level test calls
the seam object directly and therefore always can.
2. **A negative instrumentation result is worthless without a control.**
No output from an instrumented seam is only evidence once you have shown
writes from that module are visible under the harness. One
`process.stderr.write` at module load separates "never ran" from "output
swallowed" — opposite conclusions.
3. **Source-string ratchets prove syntax, not behavior.** Three were
torn down in review. The sharpest guarded a never-executed-code bug with
a source search, reproducing the bug one level up; measured, the
behavioural version fails an inverted dispatch and the textual one
passes it. Includes the sub-rules paid for the hard way: use the AST not
regex (a brace in a string truncated an extraction to 13 lines and every
count read a *passing* zero), guard the guard, anchor by index rather
than a character window.
4. **A green test on first try, on a path with no prior coverage, is a
warning.** Two conversions were reverted in one day because their tests
passed with the change reverted. Negative assertions succeed trivially
when the method returns early — `recoverCompletedTask` has seven guards
before the converted line, and the fixture has to satisfy all of them.
5. **A named workflow selection is not a resolved one.** Provenance
cannot be inferred from the returned value, because a fallback IR and a
valid id-less IR are structurally identical — the resolver that knows
has to report it. This is the fork every remaining lifecycle-column
conversion hits.
## Why this rather than another conversion
Everything left in my area is now owned and further along than I could
take it: `executor.ts` → #2628 (which solved the `recoverCompletedTask`
fixture I could not), `self-healing.ts` → #2560 (independently hit all
three traps I catalogued), the dashboard cluster → #2625/#2626/#2636.
Duplicating that would be motion, not progress. Turning findings that
cost real cycles into something greppable is the useful thing I can
still add.
`pnpm lint` clean. No changeset — internal documentation.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Documentation**
* Added a best-practices guide for verifying that workflow code paths
actually execute.
* Covers reliable behavioral assertions, instrumentation controls,
regression-proof tests, source validation, and detection of fallback
behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Six consecutive slices of U7 produced **six test-fixture defects, and
every one first presented as a bug in the code under test.** Not one was
real.
Each cost 15–60 minutes debugging the wrong file. **Two would have
shipped a false green** — a test passing while asserting nothing — if
the failure had happened to look plausible rather than implausible.
This is not a story about carelessness. Every one of these fakes was
modelled on an existing fixture in this repo, and the repo's fixtures
are inconsistent about exactly the things that matter.
## The catalogue
| # | Defect | How it presented | Real cause |
|---|---|---|---|
| 1 | `moveTaskIf` ignores its predicate | Test passed; in-txn guard
untested and indistinguishable from absent | Fake never invoked the
callback |
| 2 | `updateTaskAtomic: vi.fn()` never invokes its callback | *Every*
finalize bailed before the branch under test | Success is derived from
whether the callback ran |
| 3 | Harness default parameter swallows the input | "Task vanished"
case became a duplicate of the control | `harness(undefined)` triggers
the default |
| 4 | `logEntry: vi.fn()` returns `undefined` | Sweep appeared to match
only one column | `.catch` on a non-promise throws, aborting the loop
after item one |
| 5 | Harness lets `poll()` reach the real `specifyTask` | **exit 1 with
every test green** | Real agent path threw *asynchronously*, after
assertions passed |
| 6 | `updateTask: vi.fn()` returns `undefined` | Branch "did not run" |
Same as #4 |
**4 and 6 are the same shape, found a week apart, because nothing
prevented the second.** That is the argument for writing this down.
## The three rules
1. **Every store method a fake exposes returns what the real one
returns** — overwhelmingly a promise. Production writes `await
store.m(...).catch(h)` as a fail-soft idiom; `.catch` on `undefined`
throws a `TypeError` that unwinds into a broad *"never let housekeeping
break the poll"* handler and vanishes. Symptom is never "your fake is
wrong" — it is *"the loop only processed the first item"*.
2. **A fake handed a predicate or callback must invoke it.** Ignoring it
makes the guarded and unguarded implementations *indistinguishable*, so
a test named for the guard cannot detect the guard's removal. Includes
the `onLockedRead` hook, without which an in-transaction recheck stays
untestable even once the predicate is invoked.
3. **Stub the agent-dispatch boundary.** `poll()` ends in "start an
agent", which in a unit test throws *after* the test resolved — `17
passed`, exit code 1, which on CI reads as infrastructure noise.
> Never accept a non-zero exit on a green run. It is the only signal
that something escaped your assertions entirely.
## Also covered
- **How to spot a fixture defect fast** — the tell is *failing for the
wrong reason*. Three concrete checks before you open the production
file.
- **Why differential tests earn their keep** even when they feel
redundant: the default-vocabulary half doubles as a fixture self-check,
because it asserts behavior that is by definition already shipping. On
this program, "both halves failed" was the signal that found three of
the six.
- **The connection to guards that cannot fire** — six of those on this
program too, including a ratchet I wrote that matched only a
double-quoted literal (#2527). Same discipline either way: *prove the
check fails on the thing it claims to catch before trusting that it
passes.* Including the warning that one ratchet injection silently
failed to apply, leaving a green run that would have "proven" the
ratchet worked.
## The concrete next step, stated plainly
A shared `createTaskStoreFake({ tasks, workflowIr })` with
promise-resolving, callback-invoking defaults would remove this whole
class in one small PR. **It is not built here** because it is cross-unit
and needs adopters — building it inside U7 and hoping others find it is
how conventions die. The doc says: if you are about to hand-roll a
seventh store fake, build the helper instead and link it.
Docs-only; no changeset (AGENTS.md excludes internal docs).
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Documentation**
* Added guidance on six store-fake defect patterns that can resemble
production bugs during testing.
* Documented best practices for creating reliable store fakes, including
promise handling, callback invocation, and async dispatch isolation.
* Added diagnostic techniques for distinguishing fixture issues from
genuine application defects.
* Included guidance for validating production guards and links to
related documentation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Found while proving U11's caveat 2. **Characterization plus guard-rails
— no production change, deliberately.**
## The defect
A default-workflow card in Planning can be moved **into `triage`** — a
column its workflow no longer declares — re-creating exactly the
stranded state `reconcileUndeclaredTaskColumns` exists to repair.
Measured on a fresh store:
```
experimentalFeatures.workflowColumns null ← no production writer
createTask(...) column = "todo"
moveTask("todo" → "triage") ACCEPTED
moveTask("todo" → "bogus-column") REJECTED: "Valid targets: in-progress, triage, archived"
```
The second rejection is the tell. Validation is real — but it is the
**legacy `VALID_TRANSITIONS`** table talking, and that table does not
know the card's workflow. Its `todo` row still lists `triage`.
## Why the workflow-aware check does not run
`moves.ts` gates its adjacency block — including
`workflowHasColumn(workflowIr, toColumn)` — on
`isWorkflowColumnsCompatibilityFlagEnabled`, which reads the raw
`experimentalFeatures.workflowColumns` key. Nothing writes it, so the
block is dead on the path every real project takes.
**Corollary, already reported:** U11's undeclared-source escape hatch in
`resolveAllowedColumns` also does not run in production. It was added
with #2515 so a stranded card would have a legal move instead of `Valid
targets: none`; on the live path that rescue comes from the legacy table
instead. Mutation-verified — stubbing the hatch back to `[]` leaves the
operator-move test green.
## Why I did not fix it
PR #2499 un-gated the capacity check and **explicitly scoped validation
out**:
> SCOPE, deliberately narrow: only the CAPACITY check is un-gated.
`workflowIr` stays flag-gated so transition VALIDATION keeps its current
behavior — the inline path's bare-Error/"Valid targets:" contract is
unchanged, and none of the Phase A2 divergences are flipped here.
That is a considered decision by the owner of this function, and several
suites pin the contract it protects. Overriding it from outside would
flip an error shape I do not own.
**What has changed since that decision is U11:** the legacy table now
offers a target the default workflow does not declare, which it never
did before. That is new input to the scoping call, not licence to ignore
it — so this lands as a reproduction for U2b rather than a patch.
U2b's branch (`feature/workflow-move-path-convergence`) is stale — HEAD
predates several merged PRs, clean tree — so nothing is being raced.
## What ships
The defect is **characterized, not asserted-as-correct**: the test pins
today's behaviour so it is visible and measurable, and an `it.todo`
states the intended behaviour. Writing it as a passing "refuses" test
would have required the fix; writing it as a failing test would redden
CI; asserting the current behaviour as *correct* would be a lie.
Characterization plus `it.todo` is the honest third option.
Four guard-rails pin what a fix must **not** break:
- every declared lifecycle move (`todo → in-progress → in-review →
done`)
- archiving
- a `recoveryRehome` deliberately reaching an undeclared column — the
path that rescues already-stranded cards, and the one a careless fix
would break
- a premise test asserting the compatibility flag really is unset, so
the suite fails loudly if that ever changes rather than silently testing
a different code path
## Exposure
Narrow but real. U10 already fixed the dashboard move menu to offer only
workflow-declared targets, so the board does not present this. The
**write path** does — REST API, CLI, plugins, any stale client — which
is why the guard belongs in `moves.ts` rather than only in the UI.
5 passed + 1 todo; lint clean.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Added coverage for task moves involving workflow-declared and
undeclared columns.
* Documented a known issue where tasks can currently be moved into the
deleted `triage` column.
* Preserved valid moves, archiving, and recovery re-homing behavior.
* **Documentation**
* Added reproduction steps, affected move paths, and guardrails for
addressing the issue.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Docs only. Answers the P0 question per site: **does it still fire, what
silently stops happening, is there a backup?**
## Headline: no hard stall
The alarming reading — *"the orphaned-planning-status sweeps stop
finding default cards, so a card whose planner died sits with
`status:"planning"` forever, invisible to discovery"* — **does not
hold.**
`triage.ts`'s `sweepStalePlanningStatuses` is the **periodic primary**
for that repair and already tests `column !== "triage" && column !==
"todo"`. It covers the merged column. The two self-healing sweeps
perform the same repair and are **redundant nets**, not the sole rescue.
That is the difference between a P0 and a cleanup, and it is only
visible by reading the **backup** path rather than the broken guard.
Recorded so nobody re-derives the panic.
## Self-healing block, by blast radius
| site | fires? | what stops | backup | verdict |
|---|---|---|---|---|
| `:12106`, `:12427` | no | clearing a stale `planning` status |
`triage.sweepStalePlanningStatuses` | redundant net lost — **cleanup** |
| `:2961/2981/3016` `recoverAdvancedTriageTasks` | no | re-homing a card
with a worktree + durable IR pin to its **pinned** resume column |
hold-release still releases it on capacity (real spec ⇒
`isUnplannedForExecution` false) | **degraded, not stuck** — fix first |
| `:12254` | no | a bounded priority nudge | none needed; the doc says
nudge, not rescue | **low** |
| `:12151`, `:9151` | **yes** | — | already OR-paired | **safe** |
**Second-order trap at `:3016`.** It skips when `resumeColumn ===
"triage"`, guarding against resuming a card into the column it already
occupies. Post-merge the pinned column is `todo`, which is **not**
skipped — so pairing the literal at `:2961` *without* also pairing
`:3016` produces a `todo → todo` move. **Repair the three together.**
## Two sites in the ownership split are already handled
- **`usage-limit-detector.ts:126`** (assigned to u8) — already fixed in
**PR #2567**. Real breakage: the planning lane stopped being recognised,
so a card being planned was neither parked when its provider hit a usage
limit nor resumed when it recovered.
- **`spec-staleness.ts:95`** (assigned to u7) — already proven safe
as-is, merged with #2515. **Its obvious fix is wrong.** I tried `||
task.column === "todo"` and it turned an existing test red: it breaks
the parked-preserved-progress path.
## The generalisation, which is the most useful thing here
**On the merged column, `todo` answers two different questions.**
After the merge `todo` is both the planner column *and* the
capacity-hold column. So any site that used `triage` to mean *"is being
planned"* **cannot simply be paired with `todo`**, because `todo` also
means *"is parked waiting for capacity"*. Those sites need **status or a
trait**, not a wider literal.
That is precisely the mistake a bulk conversion makes, and
`spec-staleness.ts` is the worked example: the guard was already asking
status, and widening the column would have destroyed the distinction.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
**Merges Todo into Planning on the operator's real default workflow.**
Held from merge pending the `triage` literal audit below — see *Gating*.
## The board change
`builtin:coding` → `BUILTIN_STEPWISE_FINAL_REVIEW_CODING_WORKFLOW_IR` →
clones `BUILTIN_STEPWISE_CODING_WORKFLOW_IR`. That IR now declares
**five** columns, and `plan`, `plan-review`, `plan-replan` and `start`
all live in the merged Planning column:
```
columns: todo="Planning", in-progress, in-review, done, archived
start -> todo plan -> todo
plan-review -> todo plan-replan -> todo
parse -> in-progress (first implementation node)
```
The id stays `todo`, the display name becomes "Planning". That is the
cheaper half: `todo` was already the hold column, so every trait lookup,
task row, stored selection and the 121 `column === "todo"` guards keep
their meaning, and **no stored row needs re-homing**. Promoting `triage`
instead would have produced the same board while making those guards
workflow-*dependent* — live for Coding (Ideas), silently dead for
Coding.
`builtin:legacy-coding` keeps its six-column shape, per the operator's
decision. It exists to be the old thing.
## Entry contract, before and after each IR edit
| | result |
|---|---|
| before the default-lineage edit | **15 passed** |
| after the edit | **13 passed, 2 failed** |
| after reading both | **15 passed** |
Neither failure was routed around. One was a genuine expectation change
(two planning entry points became one); the other was my own
`mergeTodoIntoPlanning` helper throwing *"source IR is not the
split-column shape this merge transforms"* — because production **is**
the merged shape now. I **deleted** the helper rather than making it
tolerant: a transform that has silently become a no-op asserts nothing.
## The safety argument, proven not asserted
Entering at `start` is exactly what dragged cards backward in the three
earlier reverted attempts. `merged-planning-start-node-no-move.test.ts`
proves against the **real** boundary controller and **real** default IR
that entering `start` performs no move (`moveTask` is never *called*),
reaches no hold→wip capacity seam, and **still moves on a genuine
crossing** so the no-op is same-column rather than a disabled boundary.
Removing the controller's same-column short-circuit turns exactly the
two no-move tests red.
## The migration mechanism
A card can outlive its column. `resolveAllowedColumns` derives targets
from graph adjacency, and an undeclared source has none — so it returned
`[]` and **every** move was rejected with "Valid targets: none",
including the one that would rescue the card. An undeclared source now
resolves to the workflow's rebound target. Escape hatch, not relaxation:
declared columns are untouched, and it offers the rebound target *only*,
so a stranded card gets back **into** the lifecycle rather than a free
jump past review.
## A real regression this surfaced
`isDefaultWorkflowColumns` matched the legacy **six** ids as a set. The
merged default declares five, so the match stopped firing and the
default board fell through to neighbor-only adjacency, which **drops
legal moves and invents an illegal one**:
| edge | effect |
|---|---|
| `in-progress → done` | **dropped** — the mission-validation cross edge
|
| `in-review → todo` | **dropped** — review work back to planning |
| `todo/done → archived` | **dropped** — the FN-4892 direct-archival
edges |
| `done → in-review` | **invented** — a backward edge no rule allows |
Adjacency now derives from lifecycle **roles**. The load-bearing
assertion: the legacy six still reproduce `VALID_TRANSITIONS`
**verbatim**. Applied only when a workflow declares the full role set,
so custom boards keep neighbor adjacency.
## Failure accounting (core package, vs a 49-failure baseline)
| stage | failed | new |
|---|---:|---:|
| after the merge | 65 | 18 |
| after the escape hatch | 52 | 5 |
| after role-derived adjacency | 53 | 4 |
The 4 remaining are 3 `builtin-workflows` expectations encoding the
pre-merge shape and 1 create-intake expectation naming `triage` on
`builtin:coding`.
Two `schema-applier` and two `workflow-reconciliation-production-shape`
failures appeared in intermediate runs and are **not mine** — both files
pass in isolation (75/75 and 7/7). I re-ran each before attributing
them, which is why the earlier "priority" flag on the reconciliation
pair was withdrawn.
Gate: **309/309**. Lint clean.
## Gating: the `triage` audit
(`docs/solutions/architecture-patterns/u11-triage-literal-safety-audit.md`)
Program tracking cited **58** `triage` comparisons. Measured with the
same pattern:
| | count |
|---|---:|
| raw comparisons | 87 |
| inside comments | 1 |
| **not a lifecycle column at all** | **15** |
| column comparisons | 71 |
| OR-paired with `"todo"` in the same expression | 32 |
| **exclusive `triage` — the real work list** | **39** |
**15 do not compare a column.** `role === "triage"`, `surface ===
"triage"`, `sessionPurpose === "triage"`, `entry.agent === "triage"`
name the planning **agent**. Converting them would be actively wrong,
and the failure — a planning agent that can't resolve its prompt
template — would look nothing like a column bug.
**One site changes an operator-visible affordance**, which is why
per-site review beat a sweep:
`TaskCard.tsx:1927` — `taskColumnFlags?.intake === true && task.column
!== "triage"`. The literal is a **narrowing**, not a match. After the
merge a Planning card has `intake === true` and `column === "todo"`, so
the narrowing stops applying and **Start begins rendering on default
Planning cards where it previously did not.** A sweep would have
"converted" the literal and shipped the new affordance silently.
These guards do not go **dead**, they go **workflow-dependent** —
`triage` stays live for legacy-coding, Ideas, every linear built-in and
any user workflow (R11) — which is harder to detect than dead.
Work list and ownership are in the audit doc.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Docs only — no code. Companion to #2462.
##
`docs/solutions/architecture-patterns/workflow-node-column-placement-and-graph-entry-contract.md`
Why a workflow node's `column` is a lifecycle contract rather than a
display choice: it decides **who can drive the node**, **whether the
card holds a WIP slot**, and **whether anything can move it onward**.
Contents:
- The graph **entry contract** (`resolveColumnResumeNode`, shipped in
#2462) with the resume table.
- The **plan-in-place chain** — triage → finalize → continuation seed →
drain → resume → capacity suspend → release — annotated with the check
each link performs. Notably `todo`, not `triage`: an intake column has
no releaser, so a card parked there waits for a human.
- Why the pre-release gate must be narrow (column match **and**
enablement).
- The measured failure table from three reverted placement attempts.
- Why removing a column is a lifecycle-vocabulary refactor, not a
workflow edit: **82 guards** that silently stop matching, **43 writes**
to a column that no longer exists, **59 dashboard literals**. A guard
that never fires doesn't fail a test — it disables a recovery path.
## `docs/plans/2026-07-26-001-refactor-workflow-owned-lifecycle-plan.md`
The program that finishes the job, in four movements:
1. Resolve lifecycle columns from the workflow instead of ~207 string
literals.
2. Move every lane — planning, execution, review, merge — behind graph
nodes; lane services keep substrate only (storage, leases, timers,
supervision, capacity, recovery, audit).
3. A **post-commit event seam**: transitions commit transactionally,
*then* emit; subscribers react and may enqueue durable work items, but
no subscriber performs a transition. Enforced by test — dropping every
subscriber must change no lifecycle outcome.
4. Only then merge Todo into a single Planning column.
Phased so each phase lands green independently, with the IR change
deliberately **last** (KTD-7). Changing the workflow first makes the
suite green over dead guards — that's how the earlier attempts hid their
own breakage.
The merge lane **adopts** the existing design in
`docs/plans/2026-06-09-003-refactor-workflow-owned-merge-full-migration-slices-plan.md`
(slices S02–S08, still `draft-stack-handoff`) rather than authoring a
competing one, with a note to re-validate against current `main` since
it was drafted seven weeks ago.
Scale is stated honestly: ~48k lines across the four lane services, with
the executor unit explicitly landing across several commits rather than
one sweep.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Documents why self-healing force-removed a worktree a planning session was
using and parked the card branch-conflict-unrecoverable: planning gained a task
worktree but never took an active-session lease, so the reclaim sweep's liveness
guard had nothing to see, and a zero-commit branch classifies as
tip-already-merged by construction.
Captures the investigation's dead ends too — including reading maxConcurrent
from a multi-tenant config table without filtering by project_id, which produced
a confidently wrong root cause — and the three ways the first version of the fix
was itself wrong.
CONCEPTS.md: adds planning to the Active-session lease kinds (the entry had gone
stale), states the converse invariant that an unheld path reads as proof nothing
is running, and defines Top-level agent slot — the capacity concept whose
conflation with the worktree limit derailed the first hour of diagnosis.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Review of 2dbfe3d31 + 05b704dc6 surfaced real defects in both fixes:
- The unusable-worktree probe composed two helpers across an unnecessary
self-healing -> step-runner import edge, and the directory check added no
discriminating power over the `.git` probe. Replaced with one canonical
hasUsableWorktreeShape beside classifyTaskWorktree, which also applies the
repo-root gate (FN-6861) when a rootDir is available; both call sites pass one.
Its narrower guarantee vs the canonical classifier is now documented and
pinned by tests, including the de-registered shape it cannot see.
- REPLAN_PARK_STATUSES is derived from PLANNING_STAGE_STATUSES instead of
re-listed, so a new durable park status cannot be added to one set only.
- The preserve/clear decision no longer pretends to steer `worktree`: the rebound
is a reopen move, which clears it regardless. Documented, and the test now
asserts the durable row rather than only the updateTask argument.
- `branch` is cleared only when it is the re-derivable canonical fusion/<id>;
a non-canonical branch survives so a card's only commit pointer is not dropped.
- The recovery log named the recorded worktree even when the session had targeted
an AI-merge clean room. It now names the refused path and says whether the
recorded worktree was gone too.
- Added task:auto-recover-worktree-session-metadata so the decision is legible to
agents, not only in human log prose.
- isTaskStillInPlanningStage's parameter type now includes the execution stamps
its implementation reads.
- Test hygiene: real-fs fixtures wrapped in try/finally; changeset dev note
corrected; FN-8361 asserted at the discovery surface, not only in the guard
table.
Also captures the shared bug class in docs/solutions.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Global settings are already split three ways -- values in settings.json, the
revision journal in postgres, and globalMaxConcurrent/defaultProjectId in
central tables -- so the recurring "finish the cutover" proposal keeps getting
re-litigated from scratch.
Write down the two hard constraints (startup-factory reads
embeddedPostgresMaxConnections to start postgres; createFusionAuthStorage is
synchronous and host-agnostic), the recovery argument, and the one real
motivation for a partial move (multi-node policy consistency), plus the
machine-tier vs operator-policy-tier rule for placing new keys.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Review follow-up to 907e8d03e (ce-code-review, 8 personas):
- Ctrl/Cmd+V with no clipboard API (or after a denied read) now returns
false WITHOUT preventDefault: xterm skips key handling and the browser's
default paste fires xterm's helper-textarea listener once. Returning true
made non-mac xterm inject \x16 and cancel the native paste (verified
against xterm 5.5.0 _keyDown).
- A denied clipboard read sets a sticky ref so later pastes use the native
path instead of being preventDefaulted into zero delivery.
- Custom-path delivery goes through terminal.paste() to restore bracketed
paste and newline normalization.
- 'Start terminal' surfaces createTab failures in the error banner and
disables while a create is in flight (no duplicate PTY sessions).
- normalizeActiveTab() extracted so storage-read and server-validation
share one all-inactive tie-break; failure-path regression test added.
- Rewrote docs/solutions/ui-bugs/xterm-async-font-remeasure-paste-dedupe.md
to the current paste contract (was prescribing the pre-#1902 behavior).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Submit Anthropic OAuth manual codes on the first mobile tap instead of requiring keyboard dismissal first.
- Add a reusable touch action gesture hook that handles touch/pointer activation before synthetic clicks.
- Wire the OAuth manual code Submit button to invoke submission on the first touch while preventing duplicate click handling.
- Cover the mobile double-tap regression and document the UI bug pattern for future fixes.
Files changed:
.../oauth-manual-code-mobile-double-tap-submit.md | 60 +++++++++++
.../app/components/OAuthManualCodeForm.tsx | 31 +++++-
.../__tests__/OAuthManualCodeForm.test.tsx | 110 +++++++++++++++++++++
.../hooks/__tests__/useTouchActionGesture.test.ts | 110 +++++++++++++++++++++
.../dashboard/app/hooks/useTouchActionGesture.ts | 89 +++++++++++++++++
5 files changed, 399 insertions(+), 1 deletion(-)
Fusion-Task-Id: FN-7953
Fusion-Task-Lineage: d387cdbd-25a7-4b7d-add6-27a1ded5cbea
Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
Replace the boolean isGitRepository() check with a tri-state Git detection so environmental git failures (dubious ownership, missing git binary, timeouts) are no longer misreported as "not a Git repository", which previously blocked all task execution in valid repos and survived engine restarts.
- Add detectGitRepository() in worktree-pool.ts returning repo / not-repo / error (with reason: dubious-ownership, git-missing, timeout, unknown), classified from git's stderr; bound the git rev-parse call with a 10s timeout and maxBuffer; keep isGitRepository() as a backward-compatible wrapper
- Route the executor dispatch preflight guard through detectGitRepository(): only emit the original "not a Git repository / run git init" fatal on a positive not-repo verdict; on error, throw a distinct accurate error naming the real git failure, including the safe.directory remedy for dubious ownership
- Route the in-process runtime startup warning through the same tri-state detection so it only warns "not a Git repository" on a positive not-repo verdict
- Add a regression test locking extractWorktreeConflictInfo() to NOT misclassify a dubious-ownership git worktree add failure as not-git-repo
- Add targeted tests across worktree-pool, executor-worktree, and in-process-runtime test suites covering repo/not-repo/dubious-ownership/git-missing/timeout classifications on Windows OneDrive-style and POSIX paths
- Add changeset and a docs/solutions/logic-errors write-up of the false-negative root cause and fix
Files changed:
.changeset/fn-7799-git-detection-false-negative.md | 7 +++
.../logic-errors/git-detection-false-not-repo.md | 54 ++++++++++++++++
.../engine/src/__tests__/executor-worktree.test.ts | 61 +++++++++++++++++++
.../engine/src/__tests__/worktree-pool.test.ts | 71 +++++++++++++++++++---
packages/engine/src/executor.ts | 38 +++++++++---
.../runtimes/__tests__/in-process-runtime.test.ts | 53 ++++++++++++++--
packages/engine/src/runtimes/in-process-runtime.ts | 16 ++++-
packages/engine/src/worktree-pool.ts | 66 ++++++++++++++++++--
8 files changed, 334 insertions(+), 32 deletions(-)
Fusion-Task-Id: FN-7799
Fusion-Task-Lineage: 25a84283-bf47-472b-8a98-a10bf7e494de
Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
Narrative: The triage release-authorization gate itself was already removed in b5b0458; this cleans up the leftover scaffolding it left behind — an unemitted activity type, a dead TaskCard badge/label/CSS, orphaned i18n keys across all 6 locales, and a stale solutions doc — so the codebase no longer references a gate that no longer exists.
- Drop the unused `task:release-authorization-required` ActivityEventType and its label/rendering in ActivityFeed.tsx and ActivityLogModal.tsx
- Remove the dead `isReleaseAuthorizationHold` badge logic and `.awaiting-release-authorization` CSS class from TaskCard.tsx/TaskCard.css
- Simplify TaskDetailModal.tsx comments/logic now that legacy release-authorization holds render as ordinary manual plan-approval holds
- Delete orphaned i18n keys `tasks.awaitingReleaseAuthorization` and `taskDetail.plan.releaseAuthorizationHold` across en/es/fr/ko/zh-CN/zh-TW locales and resources.d.ts
- Delete the stale docs/solutions/architecture-patterns/release-triage-requires-user-authorization.md doc
- Update docs/workflow-steps.md and docs/settings-reference.md to describe the gate as removed (superseded by FN-7732) instead of documenting still-active behavior
- Add changeset for @runfusion/fusion (patch/internal)
Files changed:
.changeset/fn-7732-remove-release-authorization-block.md | 7 +++++
docs/settings-reference.md | 2 +-
docs/solutions/architecture-patterns/release-triage-requires-user-authorization.md | 33 ----------------------
docs/workflow-steps.md | 6 ++--
packages/core/src/types.ts | 8 ++++--
packages/dashboard/app/components/ActivityFeed.tsx | 5 ----
packages/dashboard/app/components/ActivityLogModal.tsx | 6 ----
packages/dashboard/app/components/TaskCard.css | 11 --------
packages/dashboard/app/components/TaskCard.tsx | 13 +++------
packages/dashboard/app/components/TaskDetailModal.tsx | 14 +++------
packages/i18n/locales/en/app.json | 3 --
packages/i18n/locales/es/app.json | 5 +---
packages/i18n/locales/fr/app.json | 5 +---
packages/i18n/locales/ko/app.json | 5 +---
packages/i18n/locales/zh-CN/app.json | 5 +---
packages/i18n/locales/zh-TW/app.json | 5 +---
packages/i18n/src/resources.d.ts | 3 --
17 files changed, 30 insertions(+), 106 deletions(-)
Fusion-Task-Id: FN-7732
Fusion-Task-Lineage: d4137bd8-9056-4062-9f2a-c6f5d47295f4
Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
Bounds durable-agent heartbeat worktree acquisition to a fixed retry count instead of requeuing to todo indefinitely across heartbeat cycles.
- Add MAX_HEARTBEAT_WORKTREE_ACQUISITION_RETRIES (3) in agent-heartbeat.ts, reusing Task.recoveryRetryCount as a cross-heartbeat counter (no schema migration)
- On cap exhaustion, terminally mark the task status:"failed" with an explanatory error, log the entry, and reopen to todo with preserveStatus so the failed status isn't wiped by reopen-to-todo semantics
- Add onTaskAcquisitionExhausted callback wired in in-process-runtime.ts to CentralCore.recordTaskCompletion(taskId, false) so exhausted acquisitions count toward totalTasksFailed
- Add regression tests in agent-heartbeat-worktree.test.ts and in-process-runtime.test.ts covering the retry cap and completion recording
- Add changeset (patch) and a docs/solutions/logic-errors writeup documenting the investigation and other worktree-collision sub-gaps found not to reproduce on HEAD
Files changed:
.changeset/fn-7721-worktree-heartbeat-retry-cap.md | 7 ++
docs/solutions/logic-errors/heartbeat-worktree-acquisition-unbounded-requeue.md | 84 ++++++++++++++++++++++
packages/engine/src/__tests__/agent-heartbeat-worktree.test.ts | 58 +++++++++++++++
packages/engine/src/__tests__/in-process-runtime.test.ts | 11 +++
packages/engine/src/agent-heartbeat.ts | 72 ++++++++++++++++++-
packages/engine/src/runtimes/in-process-runtime.ts | 12 ++++
6 files changed, 242 insertions(+), 2 deletions(-)
Fusion-Task-Id: FN-7721
Fusion-Task-Lineage: caad671c-f360-4c1c-8aaa-5b48fca5a55b
Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
Real root cause of the blank mobile terminal (the FN-7692 remeasure guard did not
fix it and is reverted here). styles.css has a mobile-only reset
`@media (max-width: 768px) { * { max-width: 100% } }` to prevent horizontal
overflow. That universal selector also matches xterm's hidden character-measurement
subtree (`.xterm-helpers` / `.xterm-char-measure-element`). That subtree's containing
block (`.xterm-helpers`) is a 0x0 absolutely-positioned box, so `max-width: 100%`
resolves to `max-width: 0` and hard-caps xterm's character-cell measurement at 0.
FitAddon.fit() then proposes 0 columns/rows and `.xterm-screen` (plus the WebGL
canvas) collapses to 0x0 — the prompt streams in and is written into xterm's row DOM
but paints into a zero-size box, so the terminal is blank. Mobile-only, which is why
desktop always rendered fine.
Reproduced live via mobile emulation: `.xterm-char-measure-element` measured 0 while
an identical monospace span in the same container measured ~295px; `max-width: none`
on the measure element restored ~295px, and reopening the terminal with the exemption
active rendered the prompt with `.xterm-screen` sized 369x760. No amount of
remeasure/refit can fix this — the CSS re-caps the measurement to 0 every time — so
the FN-7692 CharSizeService guard is removed.
- Exempt `.xterm-helpers` / `.xterm-char-measure-element` from the mobile max-width
reset in styles.css (covers both TerminalModal and SessionTerminal)
- Revert the ineffective FN-7692 remeasure guard and its tests
- Update changeset (patch) and the docs/solutions write-up to the real root cause
Note: root cause + fix validated in the automation browser via mobile emulation
(393px, iPhone UA, forced touch), not a physical device.
Fusion-Task-Id: FN-7693
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The mobile terminal rendered blank even though the WebSocket was Connected and
the shell prompt had already streamed in. Root cause (reproduced live): on the
mobile fullscreen layout xterm's CharSizeService can measure a 0-width character
cell, so FitAddon.fit() proposes 0 columns/rows and .xterm-screen (plus the WebGL
canvas) collapses to 0x0 — prompt bytes arrive and are written into xterm's row
DOM but paint into a zero-size box. Renderer-independent and mobile-layout-
specific; not fixed by resize/font-size re-fits because prior guards only validate
the container width and font load, never the resulting measured screen/cell width.
Add guardAgainstCollapsedTerminalScreen (app/utils/terminalPreferences.ts) and arm
it from both terminal surfaces (TerminalModal + SessionTerminal) right after their
initial fit. While the container has a width but .xterm-screen does not, it forces
a genuine DOM-strategy remeasure (forceTerminalFontRemeasure) + fit, re-driven by a
ResizeObserver until the screen has a real width. It waits (does not give up) while
the container is not yet measurable, is bounded so it never spins, and is disposed
on every re-init/close path. Recurrence of FN-7620/FN-7686.
- Add isTerminalScreenCollapsed + guardAgainstCollapsedTerminalScreen with tests
- Wire + dispose the guard across all xterm (re)init/close paths in both surfaces
- Add changeset (patch) and a docs/solutions write-up
Note: reproduced via mobile emulation (393px, iPhone UA, forced touch), not a
physical device; the guard is the structural fix — confirm on a real device.
Fusion-Task-Id: FN-7692
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Investigated whether --login in TerminalService first-prompt latency is a meaningful contributor and added a one-time diagnostic hint plus documentation of findings.
- Add SLOW_LOGIN_PROFILE_HINT_MS (2000ms) threshold and one-time, non-blocking console.info hint in createSession()'s PTY onData handler when a login shell is slow to produce first output
- Track spawnStartedAt and loginProfileHintLogged per session, and whether the succeeding spawn attempt used --login, without altering spawn args, timeouts, or the retry-without-login fallback
- Add regression tests covering the slow-login-profile hint behavior in terminal-service.test.ts
- Document the investigation and findings in docs/solutions/developer-experience/login-shell-profile-latency.md and link it from docs/dashboard-guide.md
- Add a patch changeset for @runfusion/fusion describing the new server-log hint
Files changed:
.changeset/fn-7688-login-shell-profile-latency.md | 7 ++
docs/dashboard-guide.md | 17 +++
.../login-shell-profile-latency.md | 79 ++++++++++++++
.../src/__tests__/terminal-service.test.ts | 121 +++++++++++++++++++++
packages/dashboard/src/terminal-service.ts | 57 ++++++++++
5 files changed, 281 insertions(+)
Fusion-Task-Id: FN-7688
Fusion-Task-Lineage: 08d5dd47-ea9f-4973-9f0e-a8d5fdeae111
Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
Speed up initial terminal load by short-circuiting the no-op server session list call when there are no persisted local tabs to validate.
- useTerminalSessions: when readTabsFromStorage returns zero tabs, skip the listTerminalSessions HTTP call entirely and mark bootstrap ready immediately, unblocking auto-create/WebSocket connect instead of serializing behind a provably-discarded round trip
- Reload-with-persisted-tabs path is unchanged and still awaits the list call since its result is decision-relevant there
- Add regression tests covering the fresh-load fast path and the persisted-tabs path
- Add changeset (patch) and a docs/solutions write-up of the bootstrap-list-serialized-before-auto-create issue
Files changed:
.changeset/fn-7686-slow-terminal-initial-load.md | 7 ++
...bootstrap-list-serialized-before-auto-create.md | 87 ++++++++++++++++++++++
.../hooks/__tests__/useTerminalSessions.test.ts | 73 ++++++++++++++++++
.../dashboard/app/hooks/useTerminalSessions.ts | 23 +++++-
4 files changed, 189 insertions(+), 1 deletion(-)
Fusion-Task-Id: FN-7686
Fusion-Task-Lineage: 9c708329-6362-4c2e-967f-aea12849c47c
Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
Root-caused and fixed the third recurrence of the mobile terminal shortcut bar not scrolling horizontally: styles.css's mobile lockdown resets touch-action to pan-y across ancestors, and touch-action's used value is the intersection of the touched element's and every ancestor's value, so the leaf .terminal-shortcut-panel's pan-x was silently defeated even though it was already correct.
- Opt the terminal overlay and modal ancestors (.modal-overlay.terminal-modal-overlay, .modal.terminal-modal--mobile, plain-media-query mobile modal, and the shortcut/status footer) into touch-action: pan-x pan-y so descendant leaf touch-action values can take effect
- Add FNXC:Terminal comments documenting the ancestor-intersection root cause and recurrence history (FN-7550/FN-7560)
- Add a documented solution note under docs/solutions/ui-bugs/ for the ancestor-intersection touch-action pattern
- Add regression tests asserting the modal/overlay/footer ancestors carry the pan-x pan-y opt-in
- Add a changeset for the fix
Files changed:
.../fn-7621-mobile-terminal-shortcut-scroll.md | 7 ++
...on-ancestor-intersection-defeats-leaf-scroll.md | 57 +++++++++++
.../dashboard/app/components/TerminalModal.css | 38 ++++++++
.../components/__tests__/TerminalModal.test.tsx | 106 +++++++++++++++++++++
4 files changed, 208 insertions(+)
Fusion-Task-Id: FN-7621
Fusion-Task-Lineage: 771fd79e-e193-43b0-908b-0e8fe2fc2c70
Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
Fixes the mobile dashboard terminal sometimes rendering completely blank on open by having TerminalModal recover from a zero/collapsed container box.
- TerminalModal now attaches a persistent ResizeObserver directly on the xterm container, mirroring SessionTerminal's existing pattern
- When the container reports a zero/collapsed box on the first post-open fit, it now re-fits once the real box settles instead of staying stuck at FitAddon's degenerate 2x1-cell floor
- Added regression tests covering TerminalModal and SessionTerminal zero-geometry recovery
- Documented the root cause and fix in docs/solutions/ui-bugs/mobile-terminal-blank-render-zero-geometry-container.md
- Added a patch changeset for @runfusion/fusion
Files changed:
.changeset/fn-7620-mobile-terminal-blank-render.md | 7 +
docs/solutions/ui-bugs/mobile-terminal-blank-render-zero-geometry-container.md | 112 +++++++
packages/dashboard/app/components/TerminalModal.tsx | 49 +++
packages/dashboard/app/components/__tests__/SessionTerminal.test.tsx | 66 ++++
packages/dashboard/app/components/__tests__/TerminalModal.test.tsx | 358 +++++++++++++++++++++
5 files changed, 592 insertions(+)
Fusion-Task-Id: FN-7620
Fusion-Task-Lineage: 103f5b17-9a6e-4e9a-ab61-65ccb2203a8d
Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
Fixes recurrence #4 of the mobile terminal excess character-spacing bug: SessionTerminal and TerminalModal now force a second genuine xterm font remeasure AFTER fitAddon.fit() settles the post-fit column count, since handleResize() never re-bakes DomRenderer's letter-spacing compensation itself.
- SessionTerminal.tsx: call forceTerminalFontRemeasure() again after fit()/sendResizeMessage() in the resize handler, re-baking spacing against the settled (post-fit) column count instead of the stale pre-fit one.
- TerminalModal.tsx: same second forceTerminalFontRemeasure() call after fitAddon.fit()/sendResize() in its resize handling path.
- Expanded SessionTerminal.test.tsx and TerminalModal.test.tsx coverage to assert the post-fit remeasure occurs.
- Added docs/solutions/ui-bugs/xterm-options-noop-remeasure-after-font-settle.md documenting recurrence #4 root cause (DomRenderer._setDefaultSpacing() never recomputes from handleResize()).
- Added changeset fn-7567-mobile-terminal-spacing.md (patch, category fix).
Files changed:
.changeset/fn-7567-mobile-terminal-spacing.md | 7 +
docs/solutions/ui-bugs/xterm-options-noop-remeasure-after-font-settle.md | 89 ++++++
packages/dashboard/app/components/SessionTerminal.tsx | 23 ++
packages/dashboard/app/components/TerminalModal.tsx | 43 ++-
packages/dashboard/app/components/__tests__/SessionTerminal.test.tsx | 127 +++++++-
packages/dashboard/app/components/__tests__/TerminalModal.test.tsx | 334 ++++++++++++++++++++-
6 files changed, 619 insertions(+), 4 deletions(-)
Fusion-Task-Id: FN-7567
Fusion-Task-Lineage: 5da20522-82d3-4c4d-9008-db71bc5b4d75
Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
Root-cause fix for xterm text rendering with wide inter-character gaps on mobile after fonts settle post-load.
- Added forceTerminalFontRemeasure() in terminalPreferences.ts to work around xterm's OptionsService setter being a no-op when reassigning an already-current fontFamily/fontSize
- Applied the remeasure helper at every post-waitForTerminalFontMetrics() settle site in TerminalModal.tsx and SessionTerminal.tsx
- Added regression tests covering the remeasure invariant across SessionTerminal, TerminalModal, and terminalPreferences
- Documented root cause and fix in docs/solutions/ui-bugs/xterm-options-noop-remeasure-after-font-settle.md
- Added changeset for @runfusion/fusion (patch)
Files changed:
.changeset/fn-7561-mobile-terminal-spacing.md | 7 +
docs/solutions/ui-bugs/xterm-options-noop-remeasure-after-font-settle.md | 120 ++++++++++
packages/dashboard/app/components/SessionTerminal.tsx | 11 +-
packages/dashboard/app/components/TerminalModal.tsx | 11 +-
packages/dashboard/app/components/__tests__/SessionTerminal.test.tsx | 108 ++++++++-
packages/dashboard/app/components/__tests__/TerminalModal.test.tsx | 257 ++++++++++++++++++++-
packages/dashboard/app/utils/__tests__/terminalPreferences.test.ts | 55 +++++
packages/dashboard/app/utils/terminalPreferences.ts | 38 +++
8 files changed, 601 insertions(+), 6 deletions(-)
Fusion-Task-Id: FN-7561
Fusion-Task-Lineage: ad9da396-abf4-46ea-8c60-f2ed40fa4b01
Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
Adds a docs/solutions learning: a Task field is silently dropped on persist
unless it has a SQLite column + defineTaskColumn + rowToTask mapping (applyTaskPatch
writes the DB-round-tripped task back over task.json). Plus a CONCEPTS.md entry for
the active-session lease (the path-keyed exclusivity/liveness registry whose key
choice caused the concurrent-workspace collision).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Document the code-review P1 as a logic-errors learning: a per-task graph
toggle (enabledWorkflowSteps, keyed by optional-group node id) collided with
the legacy step-template namespace and was silently remapped by the store
resolver, bypassing an enabled group. Cross-references the per-task-override
blast-radius cousins as the id-namespace-collision variant of that class.
Seeds an "Optional step group" entry in CONCEPTS.md.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Document the false "engine not running" banner root cause and fix as a
docs/solutions learning, and add an "Engine Singleton Lock" entry to
CONCEPTS.md: a failed per-machine lock acquisition is proof an engine is
running elsewhere, not "no engine."
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Document the test-changed shared-infra catch-all trap (gate mode replaces
affected-package coverage instead of augmenting it) under docs/solutions/,
and add an "Affected-package test selection" concept to CONCEPTS.md.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>