Files
fusion/packages/engine/src/__tests__/memory-focus-recalling.test.ts
ischindl 8fcf4bdbaa feat: Stash memory backend — session capture, per-chat backfill, opt-in vector search (#3494)
## Summary

Adds the **Stash memory backend** (`memory.backendType=stash`) that
connects Fusion's agent memory to the
[Stash](https://github.com/Fergana-Labs/stash) product — *knowledge
bases for the agent era* ([product site:
joinstash.ai](https://joinstash.ai)). Fusion becomes a first-class Stash
client: complete chat sessions and finished tasks are captured into
Stash, memory is recalled during chat, and Stash sessions are kept in
sync with the dashboard (including deletes and archival).

**Product:** <https://github.com/Fergana-Labs/stash> ·
[joinstash.ai](https://joinstash.ai)

## What's included

### 1. Stash memory backend (RUFU-068 / RUFU-121)
- New `StashMemoryBackend` (`memory.backendType=stash`) with `stashUrl`
/ `stashApiKey` settings (global secrets-store `stash-api-key` +
per-project override).
- **Complete-chat-session capture** keyed by ChatSession id.
- Sessions are classified into **per-project folders** (get-or-create,
`external_key fusion-<projectId>`, 1h per-process cache) and
**soft-deleted with their chat** via `DELETE /api/chat/sessions/:id`.
- Per-conversation **memory-focus** read-time scoping (new
`0066_chat_session_memory_focus.sql` migration — sequence renumbered
0059→0060→0061→0065→0066 as origin/main claimed the lower numbers);
event metadata enriched with `project` / `project_name` / `chat_title`.
- Recall queries normalized to single-keyword / explicit-OR ASCII (≤100
chars); shared normalizer export reused by per-turn recall.

### 2. Per-task executor transcript capture (RUFU-122)
Finished or failed tasks upload their executor transcript
(`agent-log.jsonl`) to Stash as a task session.

### 3. Bulk archive Stash sync (RUFU-125)
Archived task-planner chats soft-delete their Stash sessions on bulk
archival (paged). The snapshot of doomed session ids is taken *before*
the local bulk delete, and the Stash sync runs fire-and-forget so a
Stash stall can never delay local archival.

### 4. Per-chat "Preserve to Stash" backfill (RUFU-136)
A per-chat action that backfills a chat's full history into Stash, with
client-side idempotency and a pre-check that skips already-uploaded
content (fail-closed, no duplicate upload on transport failure).
- **Session-folder naming fix:** the first project folder is now named
"Fusion — &lt;project name&gt;" instead of the bare "Fusion" fallback
(the backfill now resolves the central-registry project name,
best-effort, never blocking the upload).

### 5. Opt-in semantic (vector) recall (RUFU-126)
`stashVectorSearch` setting (default `false` — **zero behavior change
until enabled**). For multi-word queries the backend tries Stash's
semantic-search endpoint first, then falls back byte-identically to the
keyword path. Definitive 404/405/501/503 responses are negatively cached
per process. Requires a patched Stash server (new endpoint +
`sentence-transformers` + embedding backfill); unpatched servers are
transparently bypassed after the first 404.

## Safety
- **Opt-in / inert by default:** the default backend remains `qmd`; the
Stash backend is inert until `memoryBackendType=stash` + `stashUrl` are
set.
- All Stash I/O is **best-effort, fail-closed, and non-blocking** — a
Stash outage never blocks chat, task completion, or archival. No
run-audit content is emitted.

## Testing
- Backfill + delete-sync suites (20/20), Stash backend suite (68/68),
executor memory / session capture suites, `memory-focus-recalling`,
description-guard — all green.
- `tsc` clean across core / engine / dashboard.
- Live verification: bulk backfill of 21/24 chats completed; the
"Preserve to Stash" action is idempotent on re-run.

## Changesets
- `@runfusion/fusion` **minor** — Stash memory backend + capture
(RUFU-068/121), per-task transcript (RUFU-122), bulk archive sync
(RUFU-125), per-chat backfill (RUFU-136), opt-in vector search
(RUFU-126)
- `@runfusion/fusion` **patch** — backfill session-folder naming fix


## Rebase Note (2026-08-23)

Rebased onto `origin/main` `3f448f7292` (v0.77.0-beta.7). Conflicts
resolved additively:
- `packages/core/src/postgres/schema-applier.ts` + test — upstream's
0062-0065 migrations (task/subtask splitting removal, AI merge review
reconciliation, task repository scope, FN-149 review convergence)
unioned with this PR's `chat_sessions.memory_focus` migration, which is
**renumbered 0065 → 0066** (upstream's FN-149 shipped 0065 canonically
on origin/main); `SCHEMA_BASELINE_VERSION` advances to `0066`.
- `packages/dashboard/app/components/ChatView.tsx` — upstream's docked
chat sidebar resize handlers unioned with the RUFU-136 "Preserve to
Stash" backfill handler.
- New commit: `settings.memory.*` stash-backend i18n keys added to all 6
secondary locales (RUFU-121/122 parity fix; `pnpm i18n:status` no longer
reports any violation introduced by this PR).

**Deploy note (operator environments that already ran a pre-rebase build
of this PR):** the memory-focus SQL may already be in the schema under
ledger row `0065`. Remap that row to `0066` (`UPDATE
fusion_schema_migrations SET version = '0066' WHERE version = '0065';`)
*before* first boot of a 0066-ceiling binary — otherwise the fresh
upstream `0065_fn_149_review_convergence_stage.sql` would be skipped as
"already applied". Clean databases (no prior memory-focus row) need no
action.

**CI note — Lint (lifecycle-column census) is red on the merge base:**
`pnpm check:lifecycle-columns --strict` fails identically on pure
`origin/main` `3f448f7292` with
`packages/core/src/db/legacy-adoption.ts: 0 -> 3` (3 column guards in
the U9b legacy-adoption table without a baseline entry or
`DELIBERATE-LITERAL` marker). Verified by running the census on a clean
origin/main checkout — inherited from the base, not introduced by this
PR. Fix belongs upstream; tracked separately.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added Stash memory integration with project configuration and optional
semantic search.
  * Added per-chat memory focus controls and a `/focus` command.
  * Added “Preserve to Stash” for uploading complete chat history.
  * Added automatic chat, task transcript, and completion-event capture.
* Added project-specific Stash session folders and archive/delete
synchronization.
* **Bug Fixes**
* Improved Stash folder naming and handling of missing branches during
no-commit tasks.
* **Documentation**
* Added setup, configuration, integration, vector-search, and
performance guidance.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Fusion <noreply@runfusion.ai>
Co-authored-by: gsxdsm <gsxdsm@users.noreply.github.com>
2026-08-23 16:46:14 -07:00

239 lines
10 KiB
TypeScript

import { describe, it, expect, vi, beforeEach, afterEach } from "vitest";
import { createEngineCoreMock } from "../test/mockCore.js";
import {
createMemorySearchTool,
createMemoryTools,
resolveMemorySearchTopic,
} from "../agent-tools.js";
import { buildProactiveMemoryCueBlock } from "@fusion/core";
import type { MemorySearchOptions } from "@fusion/core";
// Static namespace import so Vite statically resolves the core memory-backend
// module to the SAME instance the @fusion/core barrel (aliased to core source)
// loads internally — spying resolveMemoryBackend here honors the ESM live
// binding used by project-memory.ts (mirrors
// packages/core/src/memory/__tests__/memory-search-topic.test.ts).
import * as coreMemoryBackend from "../../../core/src/memory/memory-backend.js";
/**
* RUFU-068 engine recall-scoping tests: per-conversation memory FOCUS must
* scope project recall to a topic as a WITHIN-project read filter.
*
* The committed seam threads the topic through:
* fn_memory_search (createMemorySearchTool) -> searchProjectMemory
* -> backend.search (Stash pushes it as a &topic= search-route param, the
* SQL enforcement point, once the route supports it)
* and proactive pre-response recall (buildProactiveMemoryCueBlock) forwards
* the topic down the same path.
*
* FNXC:MemoryFocusEngineTest 2026-08-13-16:35:
* These tests run WITHOUT a live Stash / real network. They prove that
* (a) an active topic is preserved on the backend `search()` options (never a
* post-query in-memory filter), (b) undefined/null/""/"all"/"*" collapse back
* to whole-project scope (project default), (c) a whitespace-trimmed topic is
* used, and (d) focus NEVER weakens cross-project A/B isolation — the topic is
* a WITHIN-project read filter handed to the backend, and capture stays
* write-anywhere / topic-agnostic.
*/
// ── mock @fusion/core for the fn_memory_search tool wiring ─────────────
// We keep the real buildProactiveMemoryCueBlock (pass-through) and only
// spy the reads it needs; searchProjectMemory + isMemoryBackendHealthy are
// replaced with spies so createMemorySearchTool's topic threading is observed
// without any backend or network.
vi.mock("@fusion/core", async (importOriginal) => {
const actual = await importOriginal<typeof import("@fusion/core")>();
return createEngineCoreMock(
() => importOriginal<typeof import("@fusion/core")>(),
{
searchProjectMemory: vi.fn(),
isMemoryBackendHealthy: vi.fn(),
buildProactiveMemoryCueBlock: actual.buildProactiveMemoryCueBlock,
},
);
});
// The real buildProactiveMemoryCueBlock (pass-through from the mocked barrel)
// internally calls resolveMemoryBackend bound from memory-backend.js; spying the
// static namespace import above drives a fake backend through the REAL logic.
type BackendModule = typeof coreMemoryBackend;
const ROOT = "/repos/projectA";
function callTool(tool: { execute: (...args: unknown[]) => Promise<unknown> }, callId: string, params: Record<string, unknown>) {
return tool.execute(callId, params, undefined, undefined, undefined);
}
function textOf(result: unknown): string {
const first = (result as { content?: Array<{ type: string; text?: string }> })?.content?.[0];
return first?.type === "text" ? (first.text ?? "") : "";
}
describe("resolveMemorySearchTopic (RUFU-068 focus collapse)", () => {
it("returns the trimmed topic for a non-empty focus", () => {
expect(resolveMemorySearchTopic("stash lcm")).toBe("stash lcm");
expect(resolveMemorySearchTopic(" space-topics ")).toBe("space-topics");
});
it.each([undefined, null, "", " ", "all", "*", " all ", " * "])(
"collapses %j to undefined -> whole-project scope",
(focus) => {
expect(resolveMemorySearchTopic(focus as string | null | undefined)).toBeUndefined();
},
);
});
describe("fn_memory_search tool topic threading (RUFU-068)", () => {
let searchSpy: ReturnType<typeof vi.fn>;
let healthSpy: ReturnType<typeof vi.fn>;
let coreMod: typeof import("@fusion/core");
beforeEach(async () => {
coreMod = await import("@fusion/core");
searchSpy = coreMod.searchProjectMemory as unknown as ReturnType<typeof vi.fn>;
healthSpy = coreMod.isMemoryBackendHealthy as unknown as ReturnType<typeof vi.fn>;
searchSpy.mockReset();
healthSpy.mockReset();
searchSpy.mockResolvedValue([]);
healthSpy.mockResolvedValue({ backend: "stash", available: true });
});
afterEach(() => {
searchSpy.mockReset();
healthSpy.mockReset();
});
async function runSearch(opts: { params?: Record<string, unknown>; focus?: string }) {
const tool = createMemorySearchTool(ROOT, undefined, opts.focus ? { focus: opts.focus } : undefined);
return callTool(tool, "call-1", { query: "durable memory", ...opts.params });
}
it("scopes project recall to the active topic via options.focus", async () => {
await runSearch({ focus: "stash lcm" });
expect(searchSpy).toHaveBeenCalledTimes(1);
expect(searchSpy).toHaveBeenCalledWith(
ROOT,
expect.objectContaining({ query: "durable memory", limit: 5, topic: "stash lcm" }),
undefined,
);
});
it("prefers an explicit params.topic over the session focus", async () => {
await runSearch({ focus: "stash lcm", params: { topic: "explicit topic" } });
expect(searchSpy).toHaveBeenCalledWith(
ROOT,
expect.objectContaining({ query: "durable memory", limit: 5, topic: "explicit topic" }),
undefined,
);
});
it.each([undefined, "all", "*", "", " "])(
"searches whole-project scope (no topic) when focus is %j",
async (focus) => {
const tool = createMemorySearchTool(ROOT, undefined, focus !== undefined && focus !== null ? { focus } : undefined);
await callTool(tool, "call-1", { query: "durable memory" });
expect(searchSpy).toHaveBeenCalledTimes(1);
const [, options] = searchSpy.mock.calls[0] as [string, MemorySearchOptions, unknown];
expect(options.query).toBe("durable memory");
expect(options.topic).toBeUndefined();
},
);
it("threads an explicitly-set params.topic into backend.search", async () => {
await runSearch({ params: { topic: "stash lcm" } });
expect(searchSpy).toHaveBeenCalledWith(
ROOT,
expect.objectContaining({ query: "durable memory", limit: 5, topic: "stash lcm" }),
undefined,
);
});
it("is a within-project read filter: the topic reaches searchProjectMemory options, never a post-query filter", async () => {
// A topic-agnostic backend would receive topic in its options but return its
// normal (possibly unfiltered) results; the ENGINE must NOT re-filter them
// in-memory. The spy receives the topic inside search options (the seam),
// and the returned results are surfaced verbatim.
searchSpy.mockResolvedValue([
{ path: "stash://session/fusion-1-a", lineStart: 1, lineEnd: 1, snippet: "topic hit", score: 1, backend: "stash" },
{ path: "stash://session/fusion-1-b", lineStart: 1, lineEnd: 1, snippet: "unrelated hit", score: 0.5, backend: "stash" },
]);
const result = await runSearch({ focus: "stash lcm" });
// topic was forwarded to the enforcement seam (searchProjectMemory options)
expect(searchSpy.mock.calls[0][1]).toMatchObject({ topic: "stash lcm" });
// results are returned verbatim — NO client-side in-memory re-filter
expect(textOf(result)).toContain("unrelated hit");
});
it("createMemoryTools wires the same focus through fn_memory_search", async () => {
const tools = createMemoryTools(ROOT, undefined, { focus: "stash lcm" });
const searchTool = tools.find((t) => t.name === "fn_memory_search");
expect(searchTool).toBeTruthy();
await callTool(searchTool!, "call-1", { query: "durable memory" });
expect(searchSpy).toHaveBeenCalledWith(
ROOT,
expect.objectContaining({ topic: "stash lcm" }),
undefined,
);
});
});
describe("buildProactiveMemoryCueBlock topic forwarding (RUFU-068)", () => {
let backendSpy: ReturnType<typeof vi.fn>;
let resolveSpy: ReturnType<typeof vi.fn>;
beforeEach(() => {
const mod = coreMemoryBackend as unknown as BackendModule;
backendSpy = vi.fn<MemoryBackendSearch>().mockResolvedValue([]);
resolveSpy = vi.spyOn(mod, "resolveMemoryBackend").mockReturnValue({
type: "stash",
name: "Stash",
capabilities: { readable: true, writable: true, supportsAtomicWrite: false, hasConflictResolution: true, persistent: true },
read: async () => ({ content: "", exists: false, backend: "stash" }),
write: async () => ({ success: true, backend: "stash" }),
search: backendSpy,
} as never);
});
afterEach(() => {
resolveSpy.mockRestore();
});
async function cue(opts: { topic?: string }) {
return buildProactiveMemoryCueBlock(ROOT, "recall query", undefined, opts.topic !== undefined ? { topic: opts.topic } : undefined);
}
function lastSearchOptions(): MemorySearchOptions {
expect(backendSpy).toHaveBeenCalled();
const [, options] = backendSpy.mock.calls[backendSpy.mock.calls.length - 1] as [string, MemorySearchOptions];
return options;
}
it("forwards the topic to backend.search when a non-empty topic is active", async () => {
await cue({ topic: "stash lcm" });
expect(lastSearchOptions().topic).toBe("stash lcm");
});
it.each(["", "all", "*", " "])("does NOT forward a %j topic (whole-project scope)", async (topic) => {
await cue({ topic });
expect(lastSearchOptions().topic).toBeUndefined();
});
it("does not forward an omitted topic (whole-project scope)", async () => {
await cue({});
expect(backendSpy).toHaveBeenCalled();
expect(lastSearchOptions().topic).toBeUndefined();
});
it("returns results verbatim — never post-filters the recalled cue in-memory", async () => {
backendSpy.mockResolvedValue([
{ path: "stash://session/fusion-1-a", lineStart: 1, lineEnd: 1, snippet: "topic hit", score: 1, backend: "stash" },
{ path: "stash://session/fusion-1-b", lineStart: 1, lineEnd: 1, snippet: "unrelated hit", score: 0.5, backend: "stash" },
]);
const result = await cue({ topic: "stash lcm" });
expect(lastSearchOptions().topic).toBe("stash lcm");
// both results surface in the cue — the CUE is a projection of the backend's
// search result set, not a second in-memory topic filter.
expect(result).toContain("unrelated hit");
});
});
type MemoryBackendSearch = (root: string, options: MemorySearchOptions) => Promise<Array<{ path: string; lineStart: number; lineEnd: number; snippet: string; score: number; backend: string }>>;