Commit Graph

7 Commits

Author SHA1 Message Date
gsxdsm
e29954e50d Merge remote-tracking branch 'origin/main' into feature/tasks-take-too-long
# Conflicts:
#	scripts/__tests__/test-changed.test.mjs
2026-06-25 22:41:18 -07:00
gsxdsm
2cff1864c9 fix: bound @fusion/core affected lane + tighten changed watchdog under engine kill
Three changes to make `pnpm test` reliably minimal and fail gracefully:

- @fusion/core is now a memory-envelope/wide-fan-out package (was unguarded).
  It's the hub nearly everything imports (~354 test files), so a core source
  edit made `vitest --changed` expand to ~the whole core suite and blow past the
  engine's 15-min verification kill -> SIGKILL + task restart. Adding it to
  SCOPED_AFFECTED_MEMORY_ENVELOPES applies the wide-fan-out guard (run only
  directly-changed core tests, else delegate) and the bounded env. core is NOT
  gate-covered, so delegation warns loudly rather than false-greens.

- Lower CLASS_BUDGET_BANDS.changed ceiling 20min -> 13min so the script watchdog
  fails a runaway local lane itself (exit 124, no restart) BEFORE the engine's
  15-min kill restarts the whole task. A tightening, not a timeout-widening.
  Guard test pins ceiling < 900_000ms.

- Raise scoped-affected worker fan-out 1 -> 4 (operator decision). Was 1 only
  for OOM safety (FN-6854/FN-6874); the fan-out guard now bounds the set so the
  hundreds-of-files OOM driver no longer reaches these workers. Heap stays
  6144MB/worker (~4x6GB on the lane) — revisit if a RAM-constrained CI runner
  OOMs. Trades FN-5048 worker-knob guidance for throughput, scoped to the
  bounded affected lanes only.

Tests: test-changed 117/117, watchdog 15/15, eslint clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-25 21:42:22 -07:00
gsxdsm
fdccca9d27 fix(review): apply autofix feedback
- Reformat shard-floor justification as an FNXC:TestInfrastructure comment
  (project-standards: AGENTS.md FNXC_LOG convention).
- Clarify that the shard and dashboard-lane 15min floors are not coupled and
  may diverge (maintainability: avoid implying an unenforced contract).
- Add a regression-guard test pinning shard.floor=15min and asserting a 525s
  derived budget clamps up to the floor, so an accidental revert to the old
  5min floor fails loudly (correctness + testing + project-standards).
2026-06-20 21:53:31 -07:00
gsxdsm
74e056ed99 fix(ci): raise shard watchdog floor to 15min to stop false-kills
The Full Suite (non-blocking) workflow has been red for 30+ runs on main.
Diagnosis: the @fusion/engine [1/2], [2/2] and @fusion/core [2/2] shard
slices were SIGKILLed at their watchdog budgets (405s/405s/338s), not because
they hang but because those budgets are too tight for current wall-clock.

Local baselines (this machine, all pass, exit 0):
  - engine [2/2]: 145s wall / 309 files
  - core   [2/2]: 283s wall / 172 files  (old budget was only 338s!)

deriveBudgetMs tightens the budget to expected*3.5 whenever the committed
scripts/test-timings.json is <30d old. The snapshot (2026-06-03) undercounts
the import- and real-git-subprocess overhead of these heavy slices, so the
'fresh' snapshot produced a too-tight, false-kill budget on slower CI runners
-- the exact failure mode the floor/ceiling band exists to prevent.

Fix (plan KTD-2): raise the shard band floor 5min -> 15min so the heaviest
slices can't be tightened into a false-kill, while a true hang is still bounded
far under the job's 60min ceiling. Mirrors the dashboard-lane heavy-lane floor.
Follow-up: refresh scripts/test-timings.json from a default-branch CI run.
2026-06-20 21:48:14 -07:00
gsxdsm
a49450ef5e Address PR review feedback (#1669)
- watchdog: escalate forwarded SIGINT/SIGTERM/SIGHUP to SIGKILL after grace so
  external cancellation can't hang for the full budget (coderabbit major)
- watchdog: route onProcExit through signalGroup for injection consistency (greptile)
- watchdog: add cwd option; test-changed passes rootDir so pnpm runs from repo
  root regardless of invocation cwd (coderabbit major — preserved original run() cwd)
- dashboard runner: validate/clamp FUSION_RUN_VITEST_* env so a malformed value
  can't NaN-disable the watchdog (coderabbit)
- tests: verify exit-listener cleanup, forwarded-signal escalation, cwd passthrough
- plan doc: per-class-ceiling fallback wording (not median); label Output Structure fence

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-13 15:21:00 -07:00
gsxdsm
03a5d1eefb fix(review): drop redundant /* global */ comments (eslint no-redeclare)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-13 00:50:08 -07:00
gsxdsm
764abb5a47 feat(test-infra): add shared vitest watchdog + CI job timeouts
U1/U2/KTD-6: extract the dashboard heap-runner's process-group kill lifecycle
into a shared scripts/lib/run-vitest-watchdog.mjs with per-class budget bands
(timings only tighten within a generous ceiling) and an inline hang-diagnostics
snapshot. Delegate run-vitest-with-heap.mjs to it (behavior preserved). Add
timeout-minutes backstops to all full-suite.yml test jobs so a wedged run can no
longer hang to GitHub's 6h ceiling.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-13 00:32:38 -07:00