Files
fusion/packages
gsxdsm 12c4ab5a6e test(engine): pin the evaluator's archived-lane read — the service had no test at all (#3224)
## What

Pins the **evaluator's archived-lane read**. Test-only — no product
change.

`HybridEvaluatorService.evaluateTask` resolves the board's archived
lanes and hands them to `collectDeterministicSignals`, which decides
which of a task's related rows count as archived when scoring a run.

**The service had no test anywhere in the repo.** Four test files import
the module; none construct or exercise it. So this conversion was
unobservable for the simplest possible reason — *nothing ran the code*.
That is a different failure from the ones this audit has been finding
(harnesses that run the code but cannot see the difference), and worth
distinguishing: no amount of fixture care helps when the entry point is
never called.

## Measured

| | default (control) | renamed | differential |
|---|---|---|---|
| converted | pass | pass | pass |
| blinded to `["archived"]` | pass | **FAIL** | **FAIL** |

```
converted: Test Files 1 passed (1) / Tests 3 passed (3)
blinded:   Test Files 1 failed (1) / Tests 2 failed | 1 passed (3)
the 4 files importing evaluator.ts, plus this one: 5 files/76 tests, all green
lint clean; fnxc-future-dates: none added; census unchanged
```

Per the rule I documented in #3223, the blind was confirmed applied with
`git diff --stat` **before** the run rather than trusting the tool's
exit code.

## What breaks without it

On a board whose archived lane is `vaulted`, the evaluator hands the
collector the legacy `{archived}` set. Rows resting in `vaulted` are not
recognised as archived, and the deterministic half of every evaluation
score is computed from a wrong picture of the task's history. **Nothing
errors, the run completes, the number is just wrong** — which is why it
survived unnoticed.

## Pinned without faking a provider response

The assertion is about what the collector *receives*, which is decided
before any model call. `collectDeterministicSignals` is mocked to record
its arguments and throw a sentinel; the test asserts the resolved lane
set and stops.

This is deliberate over the obvious alternative of feeding `runPrompt` a
canned AI payload: `deps.runPrompt` is injectable so either approach is
offline, but a canned payload has to satisfy `parseAiResponse` and every
`EVAL_SCORE_CATEGORIES` entry, and would silently rot into a maintenance
burden on a test whose subject is one `Set`. Reversible if someone later
wants full end-to-end evaluator coverage — that is a different test, not
this one.

## Completes the engine audit

With this, every `resolveProjectColumnsForRoles` call site in
`packages/engine` has been blinded:

| file | resolvers | result |
|---|---|---|
| `self-healing.ts` | 64 | 21 pinned, 1 recorded inert by construction,
remainder mapped |
| `executor.ts` | 2 | both already covered |
| `scheduler.ts` | 1 | uncovered → pinned (#3219, merged) |
| `triage.ts` | 1 | uncovered → pinned (#3221) |
| `restart-recovery-coordinator.ts` | 1 | already covered |
| `notification-service.ts` | 1 | already covered |
| `evaluator.ts` | 1 | uncovered → pinned (this PR) |

`project-engine.ts:5154` takes `roles` as a **parameter**, so it is a
generic wrapper with no fixed role set to blind — flagged rather than
guessed at; its callers are where the question belongs.

**`packages/core`'s 17 files remain entirely unaudited** and I am
claiming nothing about them.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Tests**
* Added regression coverage to verify reliable resolution of archived
workflow lanes.
* Covered both the default archived-lane name and custom renamed
configurations.
  * Confirmed compatibility with legacy archived-lane naming behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 12:20:08 -07:00
..
2026-07-26 18:11:47 -07:00