feat(RUFU-081): add Prometheus-format /metrics system observability endpoint (#3475)

**Problem:** Fusion had no Prometheus-scrapeable metrics surface, so
production CPU/health/DB/fleet observability had to be inferred from
logs and ad-hoc monitoring.

**Fix:** Expose a Prometheus-text `/metrics` HTTP endpoint on the
dashboard with runtime + domain samplers: request latency quantiles
(p50/p95/max), last-request-age freeze indicator, CPU user/system,
RSS/heap, child-process spawn rates by kind, git children count, PG tps,
active/idle/total projects, running agents, and board task counts.
Scraped at 5s by the existing Prometheus/Grafana stack
(fusion-observability).

**Includes:** full test coverage for the Prometheus text parser and each
sampler.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added a public, unauthenticated `/metrics` endpoint with
Prometheus-formatted dashboard, runtime, process, database, and domain
metrics.
* Added request latency, resource usage, child-process, project
activity, agent, and task metrics.
* Metrics remain available in headless mode and provide stable snapshots
without triggering additional data writes or on-demand I/O.

* **Documentation**
* Documented metric families, sampling behavior, and degraded-operation
handling.

* **Tests**
* Added comprehensive coverage for collection, formatting, parsing,
endpoint behavior, and sampler lifecycle.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Fusion <noreply@runfusion.ai>
This commit is contained in:
ischindl
2026-08-19 07:11:22 +02:00
committed by GitHub
parent 027faaa09f
commit f195ff5b3d
16 changed files with 3904 additions and 0 deletions

View File

@@ -0,0 +1,7 @@
---
"@runfusion/fusion": minor
---
summary: Add a Prometheus-format /metrics observability endpoint to the dashboard.
category: feature
dev: New GET /metrics route on the dashboard server exposes runtime (process CPU user/system time, heap/RSS memory, request count and latency histogram, child-process and git-spawn counters) and domain (projects active/idle, board tasks, running agents, PostgreSQL queries per second) metric families in Prometheus text exposition format. Sampling is interval-based with an in-flight tick guard and a generation fence so a pre-close sample can never overwrite post-restart state; the process and git arms share one in-flight guard key so their coinciding default-cadence ticks skip the duplicate `ps` probe. The PostgreSQL sampler tracks counters per database: a failed-probe gap OR a dashboard stop marks the retained baseline stale (the first success after the gap/restart re-baselines and keeps the last-known rate, so a stats reset inside the gap can never produce a cross-epoch rate), and a per-database backward delta is treated as a stats reset even when the cross-database sum stays positive. Sampler start/stop are try/catch-guarded so a sampler fault can never break server startup or skip close handlers. The unauthenticated body is numeric values plus low-cardinality string labels (project identifiers and board column names are reachable to any client that can reach the port — bind to a trusted network when that is not acceptable); documented in docs/diagnostics.md. Bound to the existing dashboard port; no new network surface.

View File

@@ -249,3 +249,52 @@ No output means Git no longer registers that temp path; matching `worktree <temp
## Bounded hold-release and health diagnostics (FN-8856)
Hold-release summaries include prefetch, IR-resolution, and evaluation accumulators; measured sweep-attributable read counts; scanned-task and held-candidate counts; released/held totals; `unevaluatedCount`; and `budgetOverrunMs`. A budget-truncated sweep warns even when it stops in the preamble. Resolver reads are measured by a delegating counting facade (and direct sweep reads at their call sites), not cache-size inference. PostgreSQL health probes return a degraded timeout reason when the pool is saturated. Migration-state probe timeouts remain advisory after database and task-ID integrity checks succeed.
## `/metrics` observability endpoint (RUFU-081)
The dashboard HTTP server exposes **`GET /metrics`** — a plain-text Prometheus exposition endpoint (`text/plain; version=0.0.4; charset=utf-8`) serving the system / runtime / Fusion-domain measurements that earlier CPU and UI-responsiveness diagnoses had to collect by hand (`curl /api/health` for event-loop latency, `ps` for child-process cadence, `psql` for query rate, RAM-usage sampling for RSS). A `curl /metrics` returns the same numbers a Prometheus/Grafana scrape would consume; no OTLP collector or `prom-client` dependency is involved — the module serializes Prometheus text in-process.
### Contract
- **Public, outside `/api`:** the route is mounted at the app level (before the SPA catch-all and `express.static`), so it returns Prometheus text rather than `index.html`. Daemon bearer-token auth only protects `/api/*`; `/metrics` is intentionally unauthenticated. The body carries **numeric values plus low-cardinality string label values** — no secrets, no prose, no request payloads, no run-audit telemetry. The label values do include registered **project identifiers** (`fusion_domain_project_running_agents{project="…"}`) and board column names, so any client that can reach the port can enumerate open project ids. Bind the port to a trusted network when that disclosure is not acceptable (CodeRabbit Major review fix 2026-08-18-11:53: the earlier "numeric gauges only" wording did not match the emitted body).
- **Non-blocking by construction:** the `/metrics` handler renders **synchronously** from pre-read snapshots. It performs zero awaited I/O — a scrape completes in O(metric count) work and can never itself starve the event loop or trigger an on-demand DB/ps query. All sampling happens on pre-read tick timers (see cadence below).
- **Run-audit blackline (FN-7158/FN-7528):** no metric content, relabeled names, timestamps, or numeric snapshots are written to the run-audit. The endpoint computes on scrape from in-process state; nothing here emits a documented run-audit event.
### Served metric families
| Metric | Type | Cadence source | Meaning |
| --- | --- | --- | --- |
| `fusion_system_request_count_total` | counter | live request pipeline | Requests served through the latency recorder |
| `fusion_system_request_latency_ms{quantile="p50\|p95\|max"}` | gauge | live request pipeline | Histogram over the recent served-request ring |
| `fusion_system_request_latency_bucket{le="…"}` | gauge | live request pipeline | Cumulative bucket counts over the ring |
| `fusion_system_last_request_age_ms` | gauge | live request pipeline | ms since the last served request — **grows during event-loop starvation** (the freeze indicator) |
| `fusion_system_process_rss_bytes` | gauge | ~5s tick | RSS of the serving process |
| `fusion_system_process_heap_used_bytes` / `…_heap_total_bytes` | gauge | ~5s tick | Heap usage of the serving process |
| `fusion_system_cpu_user_seconds_total` / `…_system_seconds_total` | counter | ~5s tick | CPU time consumed by the serving process |
| `fusion_system_child_process_spawn_total` | counter | spawn hook | Cumulative `child_process` spawn/fork/execFile/exec invocations |
| `fusion_system_child_process_spawn_total_by_kind{kind="…"}` | counter | spawn hook | Per-kind cumulative spawn counts |
| `fusion_system_git_child_processes` | gauge | ~15s `ps` | Live `git` children of the serving process (best-effort) |
| `fusion_domain_postgres_queries_per_second` | gauge | ~5s tick | Derived PG xact rate from `pg_stat_database` deltas (best-effort) |
| `fusion_domain_projects_total` / `…_active` / `…_idle` | gauge | ~5s tick | Registered open project split by running-agent activity |
| `fusion_domain_project_running_agents{project="…"}` | gauge | ~5s tick | Running agents per registered project |
| `fusion_domain_board_tasks{column="…"}` | gauge | ~5s tick | Tasks per board column across registered projects |
### Sampling cadence and lifecycle
- The request-latency recorder is an Express middleware mounted before route handlers, so it measures the **live serving path** (including `GET /api/health`), not a synthetic probe.
- Process CPU/memory gauges update every ~5s; the git-subprocess gauge every ~15s; the PG rate and domain gauges every ~5s. All timers are `unref()`'d so a running sampler never keeps the process alive. The process and git arms invoke the same sampler, so they share ONE in-flight guard key: at the default 5s/15s cadence the arms coincide every 15 s and the coinciding tick skips the duplicate `ps` probe instead of double-probing (CodeRabbit Major review fix 2026-08-18-11:53).
- Sampler **start** is wired into the server's listen override and **stop** into the close handler (co-located with the OTLP exporter lifecycle), and runs in both headless and non-headless modes. Both start and stop are idempotent and never break server startup/shutdown; both calls are wrapped in try/catch (same pattern as the OTLP exporter) so a failure is logged and can never skip the remaining close handlers.
- **Spawn-count hook:** on start, `child_process.spawn`, `fork`, `execFile`, and `exec` are wrapped with an atomic counter that delegates to the original via `.apply`, so child spawning (including `superviseSpawn` / `runCommandAsync` / `execFileAsync`) is never broken. On stop the original functions are restored exactly. The hook is idempotent (starting twice never stacks a second wrap).
### Best-effort degradation
- **Git-subprocess gauge:** a single-level `ps -o comm= --ppid <pid>` scan every ~15s counts live `git` children of the serving process. It never recurses and never scans the whole process tree. When `ps` is unavailable (Windows, non-POSIX, missing procfs), the gauge degrades to `0` rather than throwing.
- **PG query-rate sampler:** reads cumulative `pg_stat_database` xact_commit/xact_rollback deltas PER DATABASE from the store's live async layer on the tick, normalized to a per-second rate. It is best-effort: on a privilege-fenced PG, transient pool error, or absent async layer it keeps the last-known rate (or `0` on the first invalid sample) rather than throwing or hammering the DB. A failed-probe gap invalidates the retained baseline — the first success after the gap re-baselines and keeps the last-known rate, so a stats reset landing inside the gap can never produce a cross-epoch rate — and a backward delta on ANY single database is treated as a stats reset even when the cross-database sum stays positive. The baseline is also marked stale on **stop**: after a dashboard stop/restart the first success re-baselines and keeps the last-known rate, so a stats reset during the stop gap can never emit a cross-epoch rate either (Greptile P1 review fix 2026-08-18-11:53). Embedded PostgreSQL reads may be operator-only depending on context.
- **Domain gauges** come only from already-open project stores via `countRunningAgentsInStore` / `listRegisteredProjectStores` / `store.listTasks({ slim: true })`. Empty/undefined/duplicate project and empty-column states produce well-formed `0`-valued or absent metric lines, never malformed output; the sampler never opens a store or starts an engine to answer a scrape.
- **Value safety:** non-finite or non-numeric values are coerced to `0` so a single bad sample cannot abort the whole body; invalid metric/label names are sanitized to the permitted Prometheus character set.
### Test coverage (RUFU-082)
- **Endpoint acceptance** (`packages/dashboard/src/routes/__tests__/metrics-endpoint.test.ts`): drives `GET /metrics` through the real server creator and an independent exposition-text parser (`packages/dashboard/src/__tests__/prometheus-text-parse.ts`) to prove the served body is well-formed Prometheus text covering all five measurement gaps and is NOT the pre-RUFU-081 SPA `index.html` fallback, that a scrape writes no run-audit row, and that repeat scrapes render a fresh, bounded snapshot.
- **Sampler acceptance** (`packages/dashboard/src/metrics/__tests__/metrics-samplers-acceptance.test.ts`): exercises the orchestrator `render()` end-to-end — synchronous pre-read render (no on-demand DB/ps on a scrape), all five family gaps as finite gauges, and the spawn-count hook incrementing on a real child process with the wrapper restored in `finally`.
- **Parser unit cases** (`packages/dashboard/src/__tests__/prometheus-text-parse.test.ts`): gauge/counter `_total`/NaN/Inf/labeled families and non-exposition-text rejection.

View File

@@ -0,0 +1,170 @@
// @vitest-environment node
import { describe, expect, it } from "vitest";
import {
assertExpositionText,
ExpositionParseError,
indexFamilies,
parseExpositionText,
requireFamily,
sampleValueOf,
unescapeLabelValue,
type ParsedMetrics,
} from "./prometheus-text-parse.js";
/**
* RUFU-082 parser unit cases.
*
* These guard the independent exposition-text grammar that the endpoint and
* sampler acceptance suites rely on, so the parser itself is covered
* first: a gauge family, a counter family with the `_total` suffix, a
* HELP/TYPE header block, a NaN sample (RUFU-081 coerces non-finite values to 0,
* but the grammar must still PARSE a literal NaN/Inf), a labeled multiline
* family, scaling/escape handling, and a non-exposition-text rejection (the
* original bug class — `GET /metrics` used to serve the SPA `index.html`).
*/
const GAUGE_BODY = [
'# HELP fusion_test_gauge A test gauge',
'# TYPE fusion_test_gauge gauge',
'fusion_test_gauge 42',
'',
].join("\n");
const COUNTER_TOTAL_BODY = [
'# HELP fusion_test_events_total Count of events',
'# TYPE fusion_test_events_total counter',
'fusion_test_events_total 7',
'',
].join("\n");
const NAN_BODY = [
'# HELP fusion_test_nan A NaN-valued gauge',
'# TYPE fusion_test_nan gauge',
'fusion_test_nan NaN',
'',
].join("\n");
const MULTILINE_BODY = [
'# HELP fusion_test_by_status Labeled gauge',
'# TYPE fusion_test_by_status gauge',
'fusion_test_by_status{status="todo"} 3',
'fusion_test_by_status{status="done"} 1',
'',
].join("\n");
describe("parseExpositionText", () => {
it("parses a scalar gauge family with its HELP/TYPE header and finite value", () => {
const parsed = parseExpositionText(GAUGE_BODY);
const family = requireFamily(parsed, "fusion_test_gauge");
expect(family.type).toBe("gauge");
expect(family.help).toBe("A test gauge");
expect(family.samples).toHaveLength(1);
expect(family.samples[0].value).toBe(42);
expect(sampleValueOf(family)).toBe(42);
});
it("parses a counter family with the _total suffix convention", () => {
const parsed = parseExpositionText(COUNTER_TOTAL_BODY);
const family = requireFamily(parsed, "fusion_test_events_total");
expect(family.type).toBe("counter");
expect(family.samples[0].value).toBe(7);
});
it("parses a literal NaN value and preserves it (grammar-level)", () => {
const parsed = parseExpositionText(NAN_BODY);
const family = requireFamily(parsed, "fusion_test_nan");
expect(Number.isNaN(family.samples[0].value)).toBe(true);
expect(family.samples[0].valueText).toBe("NaN");
});
it("parses +Inf / -Inf tokens", () => {
const body = [
'# TYPE fusion_test_inf gauge',
'fusion_test_pos_inf +Inf',
'fusion_test_neg_inf -Inf',
'',
].join("\n");
const parsed = parseExpositionText(body);
expect(requireFamily(parsed, "fusion_test_pos_inf").samples[0].value).toBe(Number.POSITIVE_INFINITY);
expect(requireFamily(parsed, "fusion_test_neg_inf").samples[0].value).toBe(Number.NEGATIVE_INFINITY);
});
it("parses a labeled multiline family into per-label samples", () => {
const parsed = parseExpositionText(MULTILINE_BODY);
const family = requireFamily(parsed, "fusion_test_by_status");
expect(family.samples).toHaveLength(2);
const byStatus = Object.fromEntries(
family.samples.map((s) => [s.labels[0]?.value, s.value]),
);
expect(byStatus).toEqual({ todo: 3, done: 1 });
});
it("treats HELP and TYPE as paired metadata blocks", () => {
const parsed = parseExpositionText(["# HELP fusion_paired help text", "# TYPE fusion_paired gauge", "fusion_paired 1", ""].join("\n"));
const family = requireFamily(parsed, "fusion_paired");
expect(family.help).toBe("help text");
expect(family.type).toBe("gauge");
expect(family.samples[0].value).toBe(1);
});
it("handles an optional epoch-millis timestamp suffix", () => {
const body = ["# TYPE fusion_test_ts gauge", "fusion_test_ts 5 1700000000123", ""].join("\n");
const parsed = parseExpositionText(body);
expect(requireFamily(parsed, "fusion_test_ts").samples[0].timestampMs).toBe(1700000000123);
});
it("ignores comment lines and blank lines", () => {
const body = ["# a comment", "", "# HELP fusion_test_c # help", "# TYPE fusion_test_c gauge", "fusion_test_c 9", "# trailing comment", ""].join("\n");
const parsed = parseExpositionText(body);
const family = requireFamily(parsed, "fusion_test_c");
expect(family.samples[0].value).toBe(9);
expect(parsed.samples).toHaveLength(1);
});
it("escapes label values per the exposition rules", () => {
const body = ['fusion_test_esc{label="a\\\"b\\\\c\\n"} 1', ""].join("\n");
const parsed = parseExpositionText(body);
const family = requireFamily(parsed, "fusion_test_esc");
expect(family.samples[0].labels[0].value).toBe('a"b\\c\n');
});
it("rejects a body that is not exposition text (HTML fallback)", () => {
const html = '<!doctype html><html><head><title>Fusion</title></head><body><div id="root"></div></body></html>';
expect(() => assertExpositionText(html)).toThrow(ExpositionParseError);
});
it("rejects a sample with a malformed numeric value", () => {
expect(() => parseExpositionText(["fusion_bad abc", ""].join("\n"))).toThrow(ExpositionParseError);
});
it("rejects an empty body", () => {
expect(() => parseExpositionText("")).toThrow(ExpositionParseError);
});
it("parses CREATE test fixtures produced by the serializer (round-trip)", () => {
// A body shaped like the endpoint's output: HELP + TYPE + value lines.
const body = [
'# HELP fusion_system_request_count_total Total HTTP requests served through the latency recorder',
'# TYPE fusion_system_request_count_total counter',
'fusion_system_request_count_total 0',
'# HELP fusion_system_process_rss_bytes Resident set size of the server process',
'# TYPE fusion_system_process_rss_bytes gauge',
'fusion_system_process_rss_bytes 1048576',
'',
].join("\n");
const parsed: ParsedMetrics = assertExpositionText(body);
const fam = indexFamilies(parsed);
expect(fam.has("fusion_system_request_count_total")).toBe(true);
expect(fam.has("fusion_system_process_rss_bytes")).toBe(true);
});
});
describe("unescapeLabelValue", () => {
it("decodes the three reserved escapes", () => {
expect(unescapeLabelValue('a\\"b\\\\c\\n')).toBe('a"b\\c\n');
});
it("leaves an unknown escape verbatim", () => {
expect(unescapeLabelValue("a\\z")).toBe("a\\z");
});
});

View File

@@ -0,0 +1,419 @@
/*
FNXC:PrometheusAcceptance 2026-08-13-16:40:
RUFU-082 acceptance tests need an INDEPENDENT Prometheus exposition-text grammar so a
scrape body is proven well-formed (and, critically, NOT the SPA index.html fallback that
GET /metrics used to serve before RUFU-081 replaced it) without sharing implementation
quirks with the serializer under test. No prometheus client library or OTLP collector is
added; this is a pure string->typed-families parser used only by tests.
*/
/**
* RUFU-082: self-contained Prometheus text exposition (version 0.0.4) parser.
*
* The `/metrics` endpoint (RUFU-081) SERIALIZES Prometheus text
* (`packages/dashboard/src/metrics/prometheus-text.ts`). This module is the
* independent acceptance-test GRAMMAR: it tokenizes a scraped body into typed
* metric families so tests can assert well-formedness, the counter `_total`
* suffix convention, NaN/Inf values, and that a body is actually parseable
* Prometheus exposition text rather than — critically — the SPA `index.html`
* fallback that `GET /metrics` used to serve before RUFU-081 replaced it.
*
* The point of a separate parser (not reusing the serializer) is that the
* acceptance test must not share implementation quirks with the thing under
* test: a serializer that produces text the same author's parser could never
* reject is a weaker guarantee than an independent grammar that does. There is
* intentionally no third-party prometheus library dependency and no network —
* this is a pure string->typed-families function.
*
* Invariants asserted here:
* - HELP/TYPE lines come in paired blocks preceding their sample lines;
* - sample lines are `name{labels} value` (timestamp optional, tolerated);
* - label values use the exposition escape rules (`\\`, `\"`, `\n`);
* - values parse as finite floats, `NaN`, `+Inf`, or `-Inf`;
* - a body that is not exposition text (e.g. an HTML shell) is rejected with
* a diff-style error naming the offending line.
*/
/** A parsed label key/value pair from a metric sample line. */
export interface ParsedLabel {
name: string;
value: string;
}
/** Type of a metric family per the exposition spec. */
export type ParsedMetricType = "counter" | "gauge" | "histogram" | "summary" | "untyped";
/** One parsed metric sample line. */
export interface ParsedSample {
/** Metric name, exactly as serialized (e.g. `fusion_system_rss_bytes`). */
name: string;
/** Label set on the sample line (empty for a scalar family). */
labels: ParsedLabel[];
/** Numeric value with NaN/±Infinity preserved from the text. */
value: number;
/** The raw numeric token (e.g. `0`, `NaN`, `+Inf`). */
valueText: string;
/** Optional epoch-millis timestamp suffix, when present. */
timestampMs?: number;
/** HELP text from the preceding `# HELP <name> ...` line, when present. */
help?: string;
/** Type from the preceding `# TYPE <name> <type>` line, when present. */
type?: ParsedMetricType;
}
/** A parsed family group: one name with its shared metadata and all samples. */
export interface ParsedFamily {
name: string;
help?: string;
type?: ParsedMetricType;
samples: ParsedSample[];
}
/** The result of a successful parse: every metric sample grouped by name. */
export interface ParsedMetrics {
families: ParsedFamily[];
/** Convenience: every sample flattened in body order. */
samples: ParsedSample[];
/** A family's `type` is implicitly `counter`/`summary` when a name ends in `_total`/`_sum`. */
}
/**
* Direct map of name -> samples; faster lookups for endpoint tests that want
* one family without scanning.
*/
export type ParsedFamilyIndex = Map<string, ParsedFamily>;
/* ------------------------------------------------------------------ *
* Tokenization primitives
* ------------------------------------------------------------------ */
const METRIC_TOKEN_RE = /^[a-zA-Z_:][a-zA-Z0-9_:]*$/;
const NUMBER_TOKEN_RE = /^[+-]?(?:(?:\d+(?:\.\d*)?)|(?:\.\d+))(?:[eE][+-]?\d+)?$/;
const LABEL_VALUE_ESCAPES: Record<string, string> = {
n: "\n",
'\\': "\\",
'"': '"',
};
/**
* Decode a Prometheus-escaped label value (`\\`, `\"`, `\n`) back to the raw
* string. Any other `\x` escape is left verbatim (spec reserves only these
* three).
*/
export function unescapeLabelValue(raw: string): string {
let out = "";
for (let i = 0; i < raw.length; i += 1) {
const ch = raw[i];
if (ch === "\\" && i + 1 < raw.length) {
const next = raw[i + 1];
const decoded = LABEL_VALUE_ESCAPES[next];
if (decoded !== undefined) {
out += decoded;
i += 1;
continue;
}
}
out += ch;
}
return out;
}
/**
* Parse the trailing `{label="value",...}` (leading brace inclusive) of a
* sample line into label pairs. Returns `[]` for a scalar line.
*/
function parseLabelSet(body: string): { labels: ParsedLabel[]; rest: string } {
if (!body.startsWith("{")) return { labels: [], rest: body };
const labels: ParsedLabel[] = [];
let i = 1;
let name = "";
// Parse `name="value"` pairs separated by commas.
while (i < body.length) {
// Skip whitespace between tokens.
while (i < body.length && /\s/.test(body[i])) i += 1;
if (body[i] === "}") {
i += 1;
break;
}
// Label name ends at `=`.
const nameStart = i;
while (i < body.length && body[i] !== "=") i += 1;
if (i >= body.length) throw new ExpositionParseError("unterminated label name");
name = body.slice(nameStart, i).trim();
// Expect `=` then `"value"`.
if (body[i] !== "=") throw new ExpositionParseError(`missing '=' after label name "${name}"`);
i += 1; // consume '='
if (body[i] !== '"') throw new ExpositionParseError(`label "${name}" value must be double-quoted`);
i += 1; // consume opening quote
let value = "";
let closed = false;
while (i < body.length) {
const ch = body[i];
if (ch === "\\") {
// Consume the escape; validate against the reserved set on decode.
if (i + 1 >= body.length) throw new ExpositionParseError("dangling escape in label value");
value += `\\${body[i + 1]}`;
i += 2;
continue;
}
if (ch === '"') {
closed = true;
i += 1;
break;
}
value += ch;
i += 1;
}
if (!closed) throw new ExpositionParseError(`unterminated value for label "${name}"`);
labels.push({ name, value: unescapeLabelValue(value) });
// Expect either `,` or `}`.
while (i < body.length && /\s/.test(body[i])) i += 1;
if (i < body.length && body[i] === ",") {
i += 1;
continue;
}
if (i < body.length && body[i] === "}") {
i += 1;
break;
}
throw new ExpositionParseError(`expected ',' or '}' after label "${name}"`);
}
return { labels, rest: body.slice(i).trim() };
}
/** Parse the numeric value token and optional timestamp of a sample line. */
function parseValueToken(token: string): { value: number; valueText: string } {
let text = token;
if (text === "NaN") return { value: Number.NaN, valueText: text };
if (text === "+Inf" || text === "Inf") return { value: Number.POSITIVE_INFINITY, valueText: text };
if (text === "-Inf") return { value: Number.NEGATIVE_INFINITY, valueText: text };
if (!NUMBER_TOKEN_RE.test(text)) {
throw new ExpositionParseError(`invalid numeric value "${token}"`);
}
return { value: Number(text), valueText: text };
}
/** An error thrown when a body is not valid Prometheus exposition text. */
export class ExpositionParseError extends Error {
readonly line: number;
readonly rawLine: string;
constructor(message: string, line?: number, rawLine?: string) {
const at = line === undefined ? "" : ` (line ${line})`;
super(`${message}${at}`);
this.name = "ExpositionParseError";
this.line = line ?? 0;
this.rawLine = rawLine ?? "";
}
}
/**
* Parse a Prometheus text exposition (version 0.0.4) body into typed metric
* families. Comments/blank lines are skipped; `# HELP` and `# TYPE` lines
* attach metadata to their following samples. Throws {@link ExpositionParseError}
* with a line-numbered diff when the body is not well-formed.
*/
export function parseExpositionText(body: string): ParsedMetrics {
const families = new Map<string, ParsedFamily>();
const samples: ParsedSample[] = [];
// Pending HELP/TYPE for a name, carried until the first sample of that name
// appears (metadata may precede samples in the body).
const pendingMeta = new Map<string, { help?: string; type?: ParsedMetricType }>();
const lines = body.split("\n");
let hadAnyLine = false;
for (let idx = 0; idx < lines.length; idx += 1) {
const lineNo = idx + 1;
const line = lines[idx];
const trimmed = line.trim();
if (trimmed.length === 0) continue;
hadAnyLine = true;
if (trimmed.startsWith("#")) {
// Comment, HELP, or TYPE directive.
const rest = trimmed.slice(1).trimStart();
if (rest.startsWith("HELP ") || rest === "HELP" || rest.startsWith("TYPE ")) {
const keyword = rest.split(/\s/, 1)[0];
const after = rest.slice(keyword.length).trimStart();
if (keyword === "TYPE") {
const parts = splitTopLevel(after);
const name = parts[0];
const typeText = parts[1];
validateNameToken(name, lineNo, line);
const type = typeText as ParsedMetricType;
if (!["counter", "gauge", "histogram", "summary", "untyped"].includes(type)) {
throw new ExpositionParseError(`unknown TYPE "${typeText}" for "${name}"`, lineNo, line);
}
const pending = getPending(pendingMeta, name);
pending.type = type;
// A TYPE line may follow an earlier HELP for the same family.
flushPending(name, pending, families, samples, lineNo, line);
} else if (keyword === "HELP") {
const name = after.split(/\s/, 1)[0];
validateNameToken(name, lineNo, line);
const help = after.slice(name.length).trim();
const pending = getPending(pendingMeta, name);
pending.help = help;
}
} else {
// A plain `# comment` or `# anything else` — valid comment line.
// Prometheus requires a space after `#` for a comment; tolerate `#foo`
// as a comment per the strictest readers? The spec says comments start
// with `# ` + a space. Treat any `#` line that isn't HELP/TYPE as a
// comment to mirror real scrape acceptance.
continue;
}
continue;
}
// Sample line: `name{labels} value[ timestamp]`.
const { name, rest } = splitMetricSample(trimmed, lineNo, line);
const { labels, rest: valueAndTs } = parseLabelSet(rest);
if (valueAndTs.length === 0) {
throw new ExpositionParseError(`sample line for "${name}" is missing a value`, lineNo, line);
}
const valueParts = valueAndTs.split(/\s+/);
const { value, valueText } = parseValueToken(valueParts[0]);
const timestampMs = valueParts.length > 1 ? Number(valueParts[1]) : undefined;
if (valueParts.length > 2) {
throw new ExpositionParseError(`sample line for "${name}" has too many tokens`, lineNo, line);
}
if (timestampMs !== undefined && !Number.isFinite(timestampMs)) {
throw new ExpositionParseError(`sample line for "${name}" has an invalid timestamp`, lineNo, line);
}
// Attach pending metadata.
const pending = pendingMeta.get(name);
const help = pending?.help;
const type = pending?.type;
pendingMeta.delete(name);
const sample: ParsedSample = {
name,
labels,
value,
valueText,
timestampMs,
help,
type,
};
samples.push(sample);
let family = families.get(name);
if (!family) {
family = { name, help, type, samples: [] };
families.set(name, family);
} else if (family.samples.length === 0) {
// First sample of this family — fill metadata from a pending/earlier line.
if (help !== undefined && family.help === undefined) family.help = help;
if (type !== undefined && family.type === undefined) family.type = type;
}
family.samples.push(sample);
}
if (!hadAnyLine) {
throw new ExpositionParseError("empty body is not valid exposition text");
}
return { families: [...families.values()], samples };
}
/** Split the metric name from the rest (`name{...} value`). */
function splitMetricSample(line: string, lineNo: number, rawLine: string): { name: string; rest: string } {
let i = 0;
while (i < line.length && !line[i].match(/[{\s]/)) i += 1;
const name = line.slice(0, i);
validateNameToken(name, lineNo, rawLine);
return { name, rest: line.slice(i).trim() };
}
/** Validate a metric or label-name token's character set. */
function validateNameToken(name: string, lineNo: number, rawLine: string): void {
if (name.length === 0 || !METRIC_TOKEN_RE.test(name)) {
throw new ExpositionParseError(`invalid metric name "${name}"`, lineNo, rawLine);
}
}
/** Get-or-create pending metadata for a family name. */
function getPending(
pendingMeta: Map<string, { help?: string; type?: ParsedMetricType }>,
name: string,
): { help?: string; type?: ParsedMetricType } {
let entry = pendingMeta.get(name);
if (!entry) {
entry = {};
pendingMeta.set(name, entry);
}
return entry;
}
/** If a family already has samples, close the pending metadata hook. */
function flushPending(
name: string,
pending: { help?: string; type?: ParsedMetricType },
families: Map<string, ParsedFamily>,
_samples: ParsedSample[],
_lineNo: number,
_rawLine: string,
): void {
const family = families.get(name);
if (!family || family.samples.length === 0) return;
if (family.help !== undefined && pending.help === undefined) pending.help = family.help;
// Metadata is consumed lazily on the next sample; nothing to flush for an
// already-populated family (its samples already captured their help/type).
void name;
}
/** Split a string on the first run of whitespace, returning both halves. */
function splitTopLevel(text: string): string[] {
const idx = text.search(/\s/);
if (idx === -1) return [text, ""];
const first = text.slice(0, idx);
const rest = text.slice(idx).trim();
return rest.length > 0 ? [first, rest] : [first, ""];
}
/* ------------------------------------------------------------------ *
* Assert helper
* ------------------------------------------------------------------ */
/**
* Assert that `body` is valid Prometheus exposition text and return its typed
* families. Throws an {@link ExpositionParseError} with a line-numbered diff
* message otherwise — small and terminal for tests.
*/
export function assertExpositionText(body: string): ParsedMetrics {
return parseExpositionText(body);
}
/** Create a name -> family index for convenient lookup in assertions. */
export function indexFamilies(parsed: ParsedMetrics): ParsedFamilyIndex {
return new Map(parsed.families.map((f) => [f.name, f]));
}
/**
* Convenience: find a family by name, asserting it exists and has at least one
* sample. Returns the single value for a scalar (no-label) single-sample family,
* otherwise the samples array.
*/
export function sampleValueOf(family: ParsedFamily | undefined): number {
if (!family || family.samples.length === 0) {
throw new ExpositionParseError("metric family has no samples");
}
if (family.samples.length !== 1) {
throw new ExpositionParseError(
`expected a scalar single-sample family, got ${family.samples.length} samples`,
);
}
return family.samples[0].value;
}
/**
* Assert (via the parser) that a metric family with the exact name exists and
* returns it. Looks through the parsed families index.
*/
export function requireFamily(parsed: ParsedMetrics, name: string): ParsedFamily {
const family = indexFamilies(parsed).get(name);
if (!family) {
throw new ExpositionParseError(`expected a metric family "${name}" but the body has none`);
}
return family;
}

View File

@@ -0,0 +1,531 @@
// @vitest-environment node
import { afterEach, describe, it, expect, vi } from "vitest";
import {
createDomainSampler,
defaultPgStatsReader,
type PgStats,
type DomainSamplerInit,
type SlimTaskLike,
} from "../domain-sampler.js";
import type { MetricFamily } from "../prometheus-text.js";
/**
* RUFU-081 domain sampler tests.
*
* Covers the injectable surface of `createDomainSampler` (mirroring the
* runtime-sampler suite):
* - the PG query-rate delta math: first sample yields 0, a real delta yields
* a per-second rate normalized by elapsed ms, and a backward/reset delta or
* an unreadable reader keeps the last-known rate (never throws);
* - the domain gauges: project total/active/idle split, per-project running
* agents, and board task counts per column, all computable from injected
* stores without a live database or engine;
* - duplicate/undefined project ids are deduped so they never produce
* malformed or duplicate metric lines;
* - `buildSnapshot` renders synchronously from pre-read state (zero awaited
* I/O), and empty/fresh states still produce well-formed 0-valued families.
*/
/** A minimal fake store with a configurable slim task list and count. */
function makeStore(tasks: SlimTaskLike[], running = 1) {
return {
listTasks: async () => tasks,
countRunning: running,
};
}
function helper<T>(
overrides: Partial<DomainSamplerInit> = {},
): ReturnType<typeof createDomainSampler> {
return createDomainSampler(overrides);
}
function sampleByName(families: MetricFamily[], name: string): MetricFamily | undefined {
return families.find((f) => f.name === name);
}
describe("PG query-rate sampler", () => {
it("first sample yields 0 (no prior delta), a later delta yields the exact per-second rate", async () => {
vi.useFakeTimers();
try {
vi.setSystemTime(Date.parse("2026-08-18T00:00:00Z"));
let stat: PgStats = [{ datname: "fusion", xactCommit: 1000, xactRollback: 10 }];
const sampler = helper({
pgStatsReader: async () => stat,
});
await sampler.samplePgRate();
expect(sampler.state.pgQueriesPerSecond).toBe(0); // first sample, no prior
// Deterministic clock: 500 commits over exactly 5 s -> exactly 100/s. An
// implementation that always returned 0 would fail the exact assertion
// (CodeRabbit Major review fix 2026-08-18-11:53: the rate tests must drive
// the clock and assert the documented invariants, not just non-NaN).
vi.setSystemTime(Date.parse("2026-08-18T00:00:05Z"));
stat = [{ datname: "fusion", xactCommit: 1500, xactRollback: 10 }];
await sampler.samplePgRate();
expect(sampler.state.pgQueriesPerSecond).toBeCloseTo(100, 10);
} finally {
vi.useRealTimers();
}
});
it("a backward / stats-reset delta keeps the last-known rate exactly (deterministic clock)", async () => {
vi.useFakeTimers();
try {
vi.setSystemTime(Date.parse("2026-08-18T00:00:00Z"));
const stats: PgStats[] = [
[{ datname: "fusion", xactCommit: 1000, xactRollback: 0 }],
[{ datname: "fusion", xactCommit: 1500, xactRollback: 0 }],
[{ datname: "fusion", xactCommit: 1400, xactRollback: 0 }], // reset: counters went backward
];
let idx = 0;
const sampler = helper({ pgStatsReader: async () => stats[Math.min(idx++, stats.length - 1)] });
await sampler.samplePgRate(); // baseline @ t0
vi.setSystemTime(Date.parse("2026-08-18T00:00:05Z"));
await sampler.samplePgRate(); // (1500-1000)/5s = exactly 100/s
const priorRate = sampler.state.pgQueriesPerSecond;
expect(priorRate).toBeCloseTo(100, 10);
vi.setSystemTime(Date.parse("2026-08-18T00:00:10Z"));
await sampler.samplePgRate(); // 1400 < 1500: per-DB backward delta = stats reset
// The reset branch preserves the prior rate EXACTLY (not just >= 0):
expect(sampler.state.pgQueriesPerSecond).toBeCloseTo(priorRate, 10);
expect(sampler.state.pgQueriesPerSecond).not.toBeNaN();
} finally {
vi.useRealTimers();
}
});
it("an unreadable PG reader degrades to last-known / 0 without throwing", async () => {
let fail = true;
const sampler = helper({
pgStatsReader: async () => {
if (fail) throw new Error("privilege-fenced");
return [{ datname: "fusion", xactCommit: 100, xactRollback: 0 }];
},
});
await expect(sampler.samplePgRate()).resolves.toBeUndefined();
expect(sampler.state.pgQueriesPerSecond).toBe(0);
fail = false;
await sampler.samplePgRate(); // baseline
expect(sampler.state.pgQueriesPerSecond).toBe(0);
fail = true;
await sampler.samplePgRate(); // now unreadable again -> keep last-known (0)
expect(sampler.state.pgQueriesPerSecond).toBe(0);
});
it("buildSnapshot emits a well-formed PG-rate family from pre-read state", async () => {
const sampler = helper({
pgStatsReader: async () => [{ datname: "fusion", xactCommit: 10, xactRollback: 2 }],
});
await sampler.samplePgRate();
const families = sampler.buildSnapshot();
const family = sampleByName(families, "fusion_domain_postgres_queries_per_second");
expect(family).toBeDefined();
expect(family!.type).toBe("gauge");
expect(family!.samples[0].value).toEqual(expect.any(Number));
expect(Number.isFinite(family!.samples[0].value)).toBe(true);
});
/*
FNXC:MetricsSampler 2026-08-16-23:35 (RUFU-081 Greptile P1 #1, RUFU-106) + 2026-08-18-04:20 review fix:
A TRANSIENT PG probe failure must never reset the sampler's baseline, AND the retained baseline
is STALE across the failed gap: a stats reset can land inside the gap (invisible to the next
probe, which may see counters already grown past the retained total), so the first success after
a failed gap RE-BASELINES and keeps the last-known rate; the NEXT success computes from the new
baseline. This deterministic fake-timer test pins both halves of the contract: an established
rate (> 0) survives a transient failure, the first post-gap sample re-baselines (no cross-gap
rate), and the following sample computes a fresh rate from the re-established baseline.
`Date.now()` is driveable via fake timers, so the expected per-second rates are exact.
*/
it("keeps the rate across a transient failure; the first post-gap sample re-baselines, the next computes fresh", async () => {
vi.useFakeTimers();
vi.setSystemTime(new Date("2026-08-16T00:00:00.000Z"));
let stat: PgStats | null = [{ datname: "fusion", xactCommit: 1000, xactRollback: 0 }];
const sampler = helper({ pgStatsReader: async () => stat });
// Baseline: first sample establishes the baseline at rate 0.
await sampler.samplePgRate();
expect(sampler.state.pgQueriesPerSecond).toBe(0);
// Advance the clock 5000ms and bump counters by 2000 -> rate 400/s.
vi.setSystemTime(new Date("2026-08-16T00:00:05.000Z"));
stat = [{ datname: "fusion", xactCommit: 3000, xactRollback: 0 }];
await sampler.samplePgRate();
expect(sampler.state.pgQueriesPerSecond).toBe(400);
// Transient failure (reader returns null) -> last-known rate MUST HOLD (400) and the
// retained baseline is marked STALE (a reset could have landed inside the gap).
vi.setSystemTime(new Date("2026-08-16T00:00:06.000Z"));
stat = null;
await sampler.samplePgRate();
expect(sampler.state.pgQueriesPerSecond).toBe(400);
// t11: first success after the gap RE-BASELINES and keeps the last-known rate; it must NOT
// emit a cross-gap rate computed from the pre-gap baseline.
vi.setSystemTime(new Date("2026-08-16T00:00:11.000Z"));
stat = [{ datname: "fusion", xactCommit: 5000, xactRollback: 0 }];
await sampler.samplePgRate();
expect(sampler.state.pgQueriesPerSecond).toBe(400);
// t16: the next success computes from the NEW baseline (5000 @ t11): (6000-5000)/5s = 200/s.
// A sampler that still used the pre-gap baseline (3000 @ t5) would report (6000-3000)/11s
// = 272.7/s instead — 200 proves the re-baseline.
vi.setSystemTime(new Date("2026-08-16T00:00:16.000Z"));
stat = [{ datname: "fusion", xactCommit: 6000, xactRollback: 0 }];
await sampler.samplePgRate();
expect(sampler.state.pgQueriesPerSecond).toBe(200);
});
/*
FNXC:MetricsSampler 2026-08-18-04:20 (RUFU-081 Greptile P1, RUFU-106 review fix):
A pg_stat_reset that lands inside a FAILED probe gap is invisible to the aggregate reset check:
the counter drops to 0 during the gap and regrows PAST the retained total before the next
successful probe, so the cross-epoch delta reads positive. The stale-gap guard must re-baseline
instead of emitting the fabricated rate.
*/
it("a stats reset inside a failed gap re-baselines instead of emitting a cross-epoch rate", async () => {
vi.useFakeTimers();
vi.setSystemTime(new Date("2026-08-16T00:00:00.000Z"));
let stat: PgStats | null = [{ datname: "fusion", xactCommit: 1000, xactRollback: 0 }];
const sampler = helper({ pgStatsReader: async () => stat });
await sampler.samplePgRate(); // baseline @ t0
vi.setSystemTime(new Date("2026-08-16T00:00:05.000Z"));
stat = [{ datname: "fusion", xactCommit: 3000, xactRollback: 0 }];
await sampler.samplePgRate(); // rate 400/s
expect(sampler.state.pgQueriesPerSecond).toBe(400);
// Failed probe @ t6 marks the baseline stale.
vi.setSystemTime(new Date("2026-08-16T00:00:06.000Z"));
stat = null;
await sampler.samplePgRate();
// During the gap a pg_stat_reset drops the counter to 0; it has now regrown PAST the
// retained total (1200 > 1000) so the cross-epoch delta (1200-1000)/5s = 40/s would read
// positive and be accepted by an aggregate-only check.
vi.setSystemTime(new Date("2026-08-16T00:00:11.000Z"));
stat = [{ datname: "fusion", xactCommit: 1200, xactRollback: 0 }];
await sampler.samplePgRate();
// The stale-gap guard re-baselines and keeps the last-known rate — no fabricated 40/s.
expect(sampler.state.pgQueriesPerSecond).toBe(400);
// The following clean sample computes from the fresh baseline: (1700-1200)/5s = 100/s.
vi.setSystemTime(new Date("2026-08-16T00:00:16.000Z"));
stat = [{ datname: "fusion", xactCommit: 1700, xactRollback: 0 }];
await sampler.samplePgRate();
expect(sampler.state.pgQueriesPerSecond).toBe(100);
});
/*
FNXC:MetricsSampler 2026-08-18-04:20 (RUFU-081 Greptile P1, RUFU-106 review fix):
A per-database stats reset can be HIDDEN by the cross-database sum: database A resets and
partially recovers while database B grows, so the aggregate delta stays positive. The per-DB
baseline check (any per-DB backward delta) is the only detector and must trigger the reset
handling (keep the last-known rate, re-baseline).
*/
it("detects a per-database stats reset that the cross-database sum would hide", async () => {
vi.useFakeTimers();
vi.setSystemTime(new Date("2026-08-16T00:00:00.000Z"));
let current: PgStats = [
{ datname: "a", xactCommit: 1000, xactRollback: 0 },
{ datname: "b", xactCommit: 0, xactRollback: 0 },
];
const sampler = helper({ pgStatsReader: async () => current });
await sampler.samplePgRate(); // baseline: a=1000, b=0
expect(sampler.state.pgQueriesPerSecond).toBe(0);
// Between samples: database a stats-reset (1000 -> 0) and recovered to 900; database b
// (no reset) grew by 200. Aggregate sum: 1000 -> 1100 — a POSITIVE delta, so an
// aggregate-only reset check would emit the fabricated (1100-1000)/5s = 20/s rate.
vi.setSystemTime(new Date("2026-08-16T00:00:05.000Z"));
current = [
{ datname: "a", xactCommit: 900, xactRollback: 0 },
{ datname: "b", xactCommit: 200, xactRollback: 0 },
];
await sampler.samplePgRate();
// Per-DB check sees a's backward delta (-100) -> reset: keep the last-known rate (0).
expect(sampler.state.pgQueriesPerSecond).toBe(0);
// And re-baselined: the next clean sample computes a real rate from the new baseline:
// (1400-900) + (700-200) = 1000 over 5s = 200/s.
vi.setSystemTime(new Date("2026-08-16T00:00:10.000Z"));
current = [
{ datname: "a", xactCommit: 1400, xactRollback: 0 },
{ datname: "b", xactCommit: 700, xactRollback: 0 },
];
await sampler.samplePgRate();
expect(sampler.state.pgQueriesPerSecond).toBe(200);
});
afterEach(() => {
vi.useRealTimers();
});
});
describe("domain gauges", () => {
it("computes project total/active/idle split and per-project running agents", async () => {
const s1 = makeStore([{ column: "todo" }], 2); // active
const s2 = makeStore([{ column: "done" }], 0); // idle
const sampler = helper({
registeredStores: () => [
{ projectId: "proj-a", store: s1 },
{ projectId: "proj-b", store: s2 },
],
countAgentsInStore: async (store: unknown) =>
(store as { countRunning: number }).countRunning,
listTasksInStore: async (store: unknown) => (store as { listTasks: () => Promise<SlimTaskLike[]> }).listTasks(),
});
await sampler.sampleDomain();
expect(sampler.state.projectCounts).toEqual({ total: 2, active: 1, idle: 1 });
expect(sampler.state.runningAgentsByProject).toEqual({ "proj-a": 2, "proj-b": 0 });
expect(sampler.state.columnCounts).toEqual({ todo: 1, done: 1 });
const families = sampler.buildSnapshot();
expect(sampleByName(families, "fusion_domain_projects_total")!.samples[0].value).toBe(2);
expect(sampleByName(families, "fusion_domain_projects_active")!.samples[0].value).toBe(1);
expect(sampleByName(families, "fusion_domain_projects_idle")!.samples[0].value).toBe(1);
expect(sampleByName(families, "fusion_domain_board_tasks")!.samples).toHaveLength(2);
});
it("dedupes duplicate / undefined project ids so metric lines are never malformed", async () => {
const store = makeStore([{ column: "todo" }], 1);
const sampler = helper({
registeredStores: () => [
{ projectId: "proj-a", store },
{ projectId: "proj-a", store }, // duplicate
{ projectId: undefined as unknown as string, store }, // malformed
],
countAgentsInStore: async () => 1,
listTasksInStore: async () => [{ column: "todo" }],
});
await sampler.sampleDomain();
expect(sampler.state.projectCounts.total).toBe(1); // only the valid, deduped id
expect(sampler.state.runningAgentsByProject).toEqual({ "proj-a": 1 });
});
it("best-effort: a throwing agent/store probe degrades to 0 without throwing and keeps a well-formed snapshot", async () => {
const sampler = helper({
registeredStores: () => [{ projectId: "proj-a", store: {} as never }],
countAgentsInStore: async () => {
throw new Error("store error");
},
listTasksInStore: async () => {
throw new Error("read error");
},
});
await expect(sampler.sampleDomain()).resolves.toBeUndefined();
expect(sampler.state.projectCounts.active).toBe(0);
const families = sampler.buildSnapshot();
// families render 0-valued (not malformed) even with an empty board
expect(sampleByName(families, "fusion_domain_board_tasks")!.samples.every((s) => s.value >= 0)).toBe(true);
});
});
describe("defaultPgStatsReader", () => {
it("returns a reader that resolves null when no registered store exposes an async layer", async () => {
const reader = defaultPgStatsReader();
// No stores registered in this test process -> resolves null (absent gauge).
await expect(reader()).resolves.toBeNull();
});
});
describe("overlap guard (RUFU-081 Greptile P1 #2)", () => {
/*
FNXC:MetricsSampler 2026-08-17-01:01 (RUFU-081 Greptile P1 #2, RUFU-106):
An async sample that outlasts its interval must never overlap the next tick of the same arm. These
fake-timer tests hold a reader's promise pending while a SECOND interval fires and assert the reader
is invoked once (the tick was skipped). Samplers therefore never run concurrently and a slow sample
never queues.
*/
it("skips a PG tick still in flight — the pg reader is invoked at most once per completed window", async () => {
vi.useFakeTimers();
let resolveReader: (s: PgStats) => void = () => {};
const pending = new Promise<PgStats>((res) => {
resolveReader = res;
});
const reader = vi.fn(() => pending);
const sampler = helper({ tick: { pgMs: 5000 }, pgStatsReader: reader });
sampler.start();
// First window fires -> reader invoked, sample stays pending (in flight).
await vi.advanceTimersByTimeAsync(5000);
expect(reader).toHaveBeenCalledTimes(1);
// Second window fires while tick 1 is still awaiting -> SKIPPED (reader NOT re-invoked).
await vi.advanceTimersByTimeAsync(5000);
expect(reader).toHaveBeenCalledTimes(1);
// Resolve the in-flight sample; the `finally` clears the guard.
resolveReader([{ datname: "fusion", xactCommit: 1000, xactRollback: 0 }]);
await vi.advanceTimersByTimeAsync(0);
// Third window: guard clear -> reader fires again, exactly once per window.
await vi.advanceTimersByTimeAsync(5000);
expect(reader).toHaveBeenCalledTimes(2);
sampler.stopTimers();
});
it("skips a domain tick still in flight — the agent/store probe is invoked at most once per completed window", async () => {
vi.useFakeTimers();
let resolveCount: (n: number) => void = () => {};
const pendingCount = new Promise<number>((res) => {
resolveCount = res;
});
const countAgents = vi.fn(() => pendingCount);
const store = makeStore([{ column: "todo" }], 1);
const sampler = helper({
tick: { domainMs: 5000 },
registeredStores: () => [{ projectId: "proj-a", store }],
countAgentsInStore: countAgents,
listTasksInStore: async () => [{ column: "todo" }],
});
sampler.start();
await vi.advanceTimersByTimeAsync(5000);
expect(countAgents).toHaveBeenCalledTimes(1);
// Second window fires while tick 1 still awaits countAgents -> SKIPPED.
await vi.advanceTimersByTimeAsync(5000);
expect(countAgents).toHaveBeenCalledTimes(1);
resolveCount(1);
await vi.advanceTimersByTimeAsync(0);
// Guard clear -> runs again.
await vi.advanceTimersByTimeAsync(5000);
expect(countAgents).toHaveBeenCalledTimes(2);
sampler.stopTimers();
});
it("still emits well-formed metric lines after the guard over an empty / non-existent store set", async () => {
vi.useFakeTimers();
// Empty registry (no registered stores) and a pg reader that resolves normally.
const sampler = helper({
tick: { pgMs: 5000, domainMs: 5000 },
registeredStores: () => [],
pgStatsReader: async () => [{ datname: "fusion", xactCommit: 10, xactRollback: 0 }],
countAgentsInStore: async () => 0,
listTasksInStore: async () => [],
});
sampler.start();
await vi.advanceTimersByTimeAsync(5000);
const families = sampler.buildSnapshot();
const pgFamily = sampleByName(families, "fusion_domain_postgres_queries_per_second");
expect(Number.isFinite(pgFamily!.samples[0].value)).toBe(true);
const projects = sampleByName(families, "fusion_domain_projects_total");
expect(projects!.samples[0].value).toBe(0);
sampler.stopTimers();
});
afterEach(() => {
vi.useRealTimers();
});
});
describe("restart fence (RUFU-081 Greptile P1 review fix)", () => {
/*
FNXC:MetricsSampler 2026-08-18-04:20 (RUFU-081 Greptile P1, RUFU-106 review fix):
A dashboard close+re-listen calls stopTimers() then start() while a pre-close sample may still
be awaiting. The fix: (a) the in-flight guard is factory-scoped so it survives the restart — a
pre-close sample still running keeps blocking ticks of the same arm until it resolves; (b) the
generation fence discards the pre-close sample's write so it can never overwrite post-restart
state. This test pins both: the pre-restart read resolves with stale data (a small counter,
which would make a leaked write baseline 100 and the next tick emit a fabricated ~225/s);
with the fence the next tick is a clean first sample (rate 0) and the following tick computes
the real (1500-1000)/5s = 100/s.
*/
it("a sample in flight at restart is fenced out and the shared in-flight guard survives the restart", async () => {
vi.useFakeTimers();
let resolveStale: (s: PgStats) => void = () => {};
const stale = new Promise<PgStats>((res) => {
resolveStale = res;
});
let calls = 0;
const reader = vi.fn(async (): Promise<PgStats | null> => {
calls += 1;
if (calls === 1) return stale; // the pre-restart tick holds this pending read
if (calls === 2) return [{ datname: "fusion", xactCommit: 1000, xactRollback: 0 }];
return [{ datname: "fusion", xactCommit: 1500, xactRollback: 0 }];
});
const sampler = helper({ tick: { pgMs: 5000, domainMs: 60_000 }, pgStatsReader: reader });
sampler.start();
// Tick 1 fires at t5000; its read is still pending when we restart.
await vi.advanceTimersByTimeAsync(5000);
expect(reader).toHaveBeenCalledTimes(1);
// Restart while tick 1 is in flight: stopTimers bumps the generation, start re-arms.
sampler.stopTimers();
sampler.start();
// The pre-restart read resolves with stale data — the fence must discard its write.
resolveStale([{ datname: "fusion", xactCommit: 100, xactRollback: 0 }]);
await vi.advanceTimersByTimeAsync(0);
// The restarted interval's first tick (t10000): without the fence the stale write would
// have baselined 100@t6000 and this tick would emit (1000-100)/4s ≈ 225/s; with the fence
// this is a clean FIRST sample -> rate 0, baseline 1000@t10000.
await vi.advanceTimersByTimeAsync(5000);
expect(reader).toHaveBeenCalledTimes(2);
expect(sampler.state.pgQueriesPerSecond).toBe(0);
// The following tick computes from the fresh post-restart baseline: (1500-1000)/5s = 100/s.
await vi.advanceTimersByTimeAsync(5000);
expect(reader).toHaveBeenCalledTimes(3);
expect(sampler.state.pgQueriesPerSecond).toBe(100);
sampler.stopTimers();
});
/*
* FNXC:MetricsSampler 2026-08-18-11:53 (RUFU-081 Greptile P1 "Restart retains stale PG
* baseline", RUFU-106 review fix): stopTimers() marks the retained PG baseline STALE, so the
* first post-restart success re-baselines and keeps the last-known rate instead of diffing
* against a pre-stop counter — otherwise a stats reset landing inside the stop gap regrows
* the counters past the retained total and the first post-restart sample emits a positive
* cross-epoch delta as a fabricated rate.
*/
it("re-baselines after a restart even when counters regrew past the stale pre-stop baseline", async () => {
vi.useFakeTimers();
vi.setSystemTime(Date.parse("2026-08-18T00:00:00Z"));
const values: PgStats[] = [
[{ datname: "fusion", xactCommit: 1000, xactRollback: 0 }],
[{ datname: "fusion", xactCommit: 1200, xactRollback: 0 }],
[{ datname: "fusion", xactCommit: 2500, xactRollback: 0 }], // regrown past the retained 1200 during the stop gap
];
let idx = 0;
const reader = vi.fn(async (): Promise<PgStats | null> => values[Math.min(idx++, values.length - 1)]);
const sampler = helper({ tick: { pgMs: 5000, domainMs: 60_000 }, pgStatsReader: reader });
sampler.start();
await vi.advanceTimersByTimeAsync(5000); // t5s: first sample, baseline 1000
await vi.advanceTimersByTimeAsync(5000); // t10s: (1200-1000)/5s = 40/s
expect(sampler.state.pgQueriesPerSecond).toBeCloseTo(40, 10);
// Restart; statistics reset during the stop gap and counters regrow to 2500 (> 1200).
sampler.stopTimers();
vi.setSystemTime(Date.parse("2026-08-18T00:01:00Z"));
sampler.start();
// First post-restart tick: the stale flag forces a re-baseline that KEEPS 40/s; without
// it the sample would emit the cross-epoch (2500-1200)/55s ≈ 23.6/s.
await vi.advanceTimersByTimeAsync(5000);
expect(reader).toHaveBeenCalledTimes(3);
expect(sampler.state.pgQueriesPerSecond).toBeCloseTo(40, 10);
sampler.stopTimers();
});
afterEach(() => {
vi.useRealTimers();
});
});

View File

@@ -0,0 +1,262 @@
// @vitest-environment node
import { EventEmitter } from "node:events";
import { describe, expect, it, vi } from "vitest";
import type { Settings, TaskStore } from "@fusion/core";
import { createServer } from "../../server.js";
import { request } from "../../test-request.js";
import {
createMetricsSampler,
type MetricsSampler,
type RuntimeSamplerInit,
} from "../index.js";
/**
* RUFU-081 endpoint integration tests.
*
* These go one level up from the unit suites: they mount the real
* `/metrics` route on a server assembled via `createServer` (the same flow the
* `test-request.ts` harness drives for other app-level routes) and assert the
* actual Prometheus-text body a `curl /metrics` would receive. They also cover
* the orchestrator's lifecycle contract that `server.ts` depends on (start/stop
* idempotency, spawn-hook install/removal).
*
* Surfaces asserted here (from the task's Surface Enumeration):
* - the `/metrics` route returns `text/plain; version=0.0.4` and never falls
* through to the SPA shell;
* - a nonexistent route still reproduces the pre-metrics behavior (SPA shell
* in non-headless, default 404 in headless);
* - headless mode still mounts the route (API/websocket-only servers expose
* the same scrape surface);
* - sampler start/stop is idempotent and stop removes the spawn hook;
* - the render path is synchronous (zero awaited I/O in a scrape).
*/
/** Minimal store double (mirrors `task-effective-settings-route.test.ts`). */
class MockStore extends EventEmitter {
getRootDir(): string { return "/repo"; }
getFusionDir(): string { return "/repo/.fusion"; }
// FNXC:PostgresCutover: server setup probes the async layer, so the route
// double exposes the production-shaped backend seam.
getAsyncLayer = vi.fn(() => ({
db: {
update: vi.fn(() => ({
set: vi.fn(() => ({
where: vi.fn(() => ({ returning: vi.fn(async () => []) })),
})),
})),
},
}));
getSettings = vi.fn(async () => this.getSettingsFast());
getSettingsFast = vi.fn(async (): Promise<Settings> => ({} as Settings));
getTaskWorkflowSelection = vi.fn(() => undefined);
getWorkflowDefinition = vi.fn(async () => undefined);
getWorkflowSettingValues = vi.fn(() => ({}));
getWorkflowSettingsProjectId = vi.fn(() => "default");
getProjectScopedPluginMcpServers = vi.fn().mockResolvedValue([]);
}
/** Build a server app (non-headless by default). */
function createApp(opts: { headless?: boolean } = {}) {
return createServer(new MockStore() as unknown as TaskStore, {
noAuth: true,
headless: opts.headless,
});
}
/** A fake-timer-friendly interval surface. */
function fakeTimers() {
const intervals = new Set<{ unref?: () => void }>();
return {
setInterval: (fn: () => void, _ms: number) => {
const token = { unref: () => undefined };
intervals.add(token);
void fn;
return token;
},
clearInterval: (t: { unref?: () => void }) => {
intervals.delete(t);
},
/** Number of currently-active fake intervals. */
activeCount: () => intervals.size,
};
}
/** A no-op process + spawn + ps surface so the orchestrator is inert for lifecycle tests. */
function inertInit(): {
runtime: { processRef: RuntimeSamplerInit["processRef"]; psProbe: RuntimeSamplerInit["psProbe"] };
} {
return {
runtime: {
processRef: {
pid: 1234,
cpuUsage: () => ({ user: 100, system: 50 }),
memoryUsage: () => ({ rss: 1_000_000, heapTotal: 512_000, heapUsed: 256_000 }),
},
psProbe: async () => ({ ok: true, childCommands: [] }),
},
};
}
describe("GET /metrics (app-level route)", () => {
it("serves Prometheus text (text/plain; version=0.0.4) with expected HELP/TYPE/sample lines", async () => {
const app = createApp();
const res = await request(app, "GET", "/metrics");
expect(res.status).toBe(200);
const contentType = String(res.headers["content-type"] ?? "");
expect(contentType).toContain("text/plain");
expect(contentType).toContain("version=0.0.4");
const body = String(res.body);
// Never the SPA shell.
expect(body).not.toContain("<!doctype html>");
expect(body).not.toContain("<html");
expect(body).not.toContain("id=\"root\"");
// Core runtime metrics present with proper HELP/TYPE exposition.
expect(body).toContain("# HELP fusion_system_request_count_total");
expect(body).toContain("# TYPE fusion_system_request_count_total counter");
expect(body).toContain("fusion_system_request_count_total 0\n");
// Request-latency quantile family (labeled; the ring is empty pre-request).
expect(body).toContain("# TYPE fusion_system_request_latency_ms gauge");
expect(body).toContain('fusion_system_request_latency_ms{quantile="p50"} 0');
expect(body).toContain('fusion_system_request_latency_ms{quantile="p95"} 0');
expect(body).toContain('fusion_system_request_latency_ms{quantile="max"} 0');
// Last-request-age gauge (the freeze indicator).
expect(body).toContain("# TYPE fusion_system_last_request_age_ms gauge");
expect(body).toMatch(/^fusion_system_last_request_age_ms 0$/m);
// Process gauges are finite and present.
expect(body).toContain("# TYPE fusion_system_process_rss_bytes gauge");
expect(body).toMatch(/^fusion_system_process_rss_bytes \d+$/m);
// Spawn counter scalar is present (0 before any spawn).
expect(body).toContain("# TYPE fusion_system_child_process_spawn_total counter");
expect(body).toContain("fusion_system_child_process_spawn_total 0\n");
// Git gauge present (best-effort 0).
expect(body).toContain("# TYPE fusion_system_git_child_processes gauge");
expect(body).toContain("fusion_system_git_child_processes 0\n");
// Domain gauges present (empty registrar -> well-formed 0-valued lines).
expect(body).toContain("# TYPE fusion_domain_postgres_queries_per_second gauge");
expect(body).toContain("fusion_domain_postgres_queries_per_second 0\n");
expect(body).toContain("# TYPE fusion_domain_projects_total gauge");
expect(body).toContain("fusion_domain_projects_total 0\n");
});
it("records a served request through the real pipeline and the latency quantiles reflect it", async () => {
const app = createApp();
// GET /api/health goes through the real request pipeline + latency recorder.
const health = await request(app, "GET", "/api/health");
expect(health.status).toBe(200);
// The scrape renders SYNCHRONOUSLY from pre-read state, so it can only see
// requests completed before it (the health request). A scrape can never
// count itself (its own `finish` fires after render).
const res = await request(app, "GET", "/metrics");
expect(res.status).toBe(200);
const body = String(res.body);
// The health request was recorded through the real pipeline.
const countMatch = body.match(/^fusion_system_request_count_total (\d+)$/m);
expect(countMatch).not.toBeNull();
expect(Number(countMatch![1])).toBeGreaterThanOrEqual(1);
// The age gauge is small (< 5s) because a request was just served.
const ageMatch = body.match(/^fusion_system_last_request_age_ms (\d+)$/m);
expect(ageMatch).not.toBeNull();
expect(Number(ageMatch![1])).toBeLessThan(5000);
});
it("is available in headless mode too (API/websocket-only server)", async () => {
const app = createApp({ headless: true });
const res = await request(app, "GET", "/metrics");
expect(res.status).toBe(200);
const body = String(res.body);
expect(body).toContain("# TYPE fusion_system_request_count_total counter");
});
it("a nonexistent route still reproduces prior behavior (SPA shell in non-headless, 404 in headless)", async () => {
// Non-headless: the SPA fallback serves the index shell for navigation paths.
const app = createApp();
const spaRes = await request(app, "GET", "/some/nonexistent/path");
// The SPA fallback serves index.html OR a 404 for file-like paths; either is
// acceptable — what matters is that /some/nonexistent/path is NOT served by
// the /metrics handler (returning Prometheus text).
const spaBody = String(spaRes.body);
expect(spaBody).not.toContain("# HELP fusion_system_request_count_total");
// Headless: default express 404 (no SPA shell, no metrics body).
const headlessApp = createApp({ headless: true });
const headlessRes = await request(headlessApp, "GET", "/some/nonexistent/path");
expect(headlessRes.status).toBe(404);
expect(String(headlessRes.body)).not.toContain("# HELP fusion_system_request_count_total");
});
});
describe("metrics sampler orchestrator lifecycle", () => {
it("start()/stop() is idempotent and stop() removes the spawn hook exactly", () => {
// A spawn module we can inspect for wrapping.
const originalSpawn = vi.fn(() => ({ on: vi.fn(), kill: vi.fn() }));
const originalFork = vi.fn(() => ({}));
const originalExecFile = vi.fn(() => ({}));
const originalExec = vi.fn(() => ({}));
const spawnMod = {
spawn: originalSpawn,
fork: originalFork,
execFile: originalExecFile,
exec: originalExec,
} as unknown as RuntimeSamplerInit["spawnModule"];
const timers = fakeTimers();
const sampler: MetricsSampler = createMetricsSampler({
runtime: {
...inertInit().runtime,
spawnModule: spawnMod,
timers,
},
domain: { timers },
});
expect(timers.activeCount()).toBe(0);
sampler.start();
expect(sampler.started).toBe(true);
expect(sampler.runtime.spawnHookInstalled).toBe(true);
// start() is idempotent: a second start never stacks a second wrap.
sampler.start();
expect(sampler.started).toBe(true);
// The spawn functions are wrapped once (not the original after install).
expect(spawnMod.spawn).not.toBe(originalSpawn);
expect(spawnMod.fork).not.toBe(originalFork);
// Timers armed.
expect(timers.activeCount()).toBeGreaterThan(0);
sampler.stop();
expect(sampler.started).toBe(false);
expect(sampler.runtime.spawnHookInstalled).toBe(false);
// The originals were restored EXACTLY on stop.
expect(spawnMod.spawn).toBe(originalSpawn);
expect(spawnMod.fork).toBe(originalFork);
expect(spawnMod.execFile).toBe(originalExecFile);
expect(spawnMod.exec).toBe(originalExec);
// All timers cleared.
expect(timers.activeCount()).toBe(0);
// stop() is idempotent (no throw, no re-wrap).
sampler.stop();
expect(spawnMod.spawn).toBe(originalSpawn);
});
it("render() is synchronous and produces a deterministic body from pre-read state", () => {
const sampler = createMetricsSampler({ runtime: inertInit().runtime });
const a = sampler.render(1_700_000_000_000);
const b = sampler.render(1_700_000_000_000);
expect(a).toBe(b);
expect(a.length).toBeGreaterThan(0);
expect(a).toContain("# TYPE fusion_system_request_count_total counter");
});
});

View File

@@ -0,0 +1,254 @@
// @vitest-environment node
import { createRequire } from "node:module";
import { describe, expect, it, vi } from "vitest";
import {
assertExpositionText,
requireFamily,
sampleValueOf,
type ParsedMetrics,
} from "../../__tests__/prometheus-text-parse.js";
import {
createMetricsSampler,
createRuntimeSampler,
type MetricsSampler,
type RuntimeSampler,
type RuntimeSamplerInit,
} from "../index.js";
/*
FNXC:PrometheusAcceptance 2026-08-13-16:52:
RUFU-082 sampler acceptance: built on top of RUFU-081's module-level unit suites,
this exercises the ORCHESTRATOR render() end-to-end through the independent
parser to prove (a) a scrape is a synchronous pre-read snapshot that never runs
an on-demand DB probe, (b) the full body covers all five measurement gaps, and
(c) the spawn-count hook increments on a real child process and is restored in
finally so no test poisons sibling spawns. The samplers live in
packages/dashboard/src/metrics/, so the suite is co-located there rather than in
packages/core/src/process (the PROMPT's assumption predates RUFU-081's layout).
*/
/**
* RUFU-082 sampler acceptance suite, exercising the ORCHESTRATOR `render()`
* end-to-end through the independent exposition parser.
*
* RUFU-081 shipped unit suites for each sampler module (they assert `buildSnapshot`
* arrays in isolation). This suite goes one level up: it renders the FULL
* Prometheus text body via `createMetricsSampler().render()` and asserts:
* - render is a SYNCHRONOUS pre-read snapshot — a scrape never runs a
* blocking/on-demand DB probe or ps scan (the pre-read gauge-bookkeeping
* seam the task requires);
* - the rendered body is well-formed exposition text covering all five
* measurement gaps;
* - a real child-process spawn increments the spawn counter hook on the LIVE
* node:child_process module and the child actually ran, with the wrapper
* restored in `finally` so no test poisons sibling spawns.
*
* No production DB queries, no polling, no real network, no time waits.
*/
/** Inert runtime init: no-op process + ps surface, no timers. */
function inertRuntime(): RuntimeSamplerInit {
return {
processRef: {
pid: 1234,
cpuUsage: () => ({ user: 1_000_000, system: 500_000 }),
memoryUsage: () => ({ rss: 2_000_000, heapTotal: 1_000_000, heapUsed: 400_000 }),
},
psProbe: async () => ({ ok: true, childCommands: ["git"] }),
};
}
/** Run one orchestrator render and parse it. */
function renderAndParse(sampler: MetricsSampler): ParsedMetrics {
return assertExpositionText(sampler.render());
}
const FIVE_GAP_FAMILIES = [
"fusion_system_request_latency_ms",
"fusion_system_last_request_age_ms",
"fusion_system_child_process_spawn_total",
"fusion_domain_postgres_queries_per_second",
"fusion_system_git_child_processes",
"fusion_system_cpu_user_seconds_total",
"fusion_system_process_rss_bytes",
];
describe("orchestrator render() acceptance", () => {
it("render() is a synchronous pre-read snapshot and never touches the DB/ps on a scrape", () => {
// A domain PG reader that THROWS if a scrape tried to reach it. The
// acceptance contract is: measurement cost is paid on the tick, not on the
// render/scrape. render() must render purely from pre-read state.
const pgProbe = vi.fn(async () => {
throw new Error("render must not issue an on-demand DB probe");
});
const sampler = createMetricsSampler({
runtime: inertRuntime(),
domain: { pgStatsReader: pgProbe },
});
// render() does not await and returns synchronously.
const body = sampler.render();
expect(typeof body).toBe("string");
expect(body.length).toBeGreaterThan(0);
expect(pgProbe).not.toHaveBeenCalled();
// Every scalar family value is a finite number (the serializer coerces NaN).
const parsed = assertExpositionText(body);
for (const sample of parsed.samples) {
expect(typeof sample.value).toBe("number");
}
});
it("renders a well-formed body covering all five RUFU-081 measurement gaps", () => {
const sampler = createMetricsSampler({ runtime: inertRuntime() });
const parsed = renderAndParse(sampler);
for (const familyName of FIVE_GAP_FAMILIES) {
const family = requireFamily(parsed, familyName);
// A family must have samples (scalar == 1, quantile-labeled == 3) and
// every value must be a finite number (the serializer coerces NaN/Inf
// to 0 so a ragged gauge never produces an unparseable body).
expect(family.samples.length).toBeGreaterThan(0);
for (const sample of family.samples) {
expect(Number.isFinite(sample.value)).toBe(true);
}
}
});
it("renders a stable, bounded family set across repeated renders (pre-read bookkeeping)", () => {
const sampler = createMetricsSampler({ runtime: inertRuntime() });
const first = renderAndParse(sampler);
const second = renderAndParse(sampler);
const shape = (p: ParsedMetrics) => p.families.map((f) => [f.name, f.samples.length]);
expect(shape(second)).toEqual(shape(first));
});
it("reflects a request recorded through the latency recording head in the rendered body", () => {
const sampler = createMetricsSampler({ runtime: inertRuntime() });
// The head is what server.ts mounts before route handlers; driving it with
// a fake response whose `finish` we fire reproduces a live served request.
const state = sampler.runtime.latency;
const middleware = sampler.middleware();
const finishCb = { once: (_ev: string, cb: () => void) => cb() } as unknown as Record<string, unknown>;
let nextCalled = false;
middleware(null, finishCb, () => {
nextCalled = true;
});
expect(nextCalled).toBe(true);
const body = sampler.render();
const parsed = assertExpositionText(body);
expect(requireFamily(parsed, "fusion_system_request_count_total").samples[0].value).toBeGreaterThanOrEqual(1);
expect(state.requestCount).toBeGreaterThanOrEqual(1);
// The freeze indicator reflects the just-recorded request (small, finite).
const age = sampleValueOf(requireFamily(parsed, "fusion_system_last_request_age_ms"));
expect(Number.isFinite(age)).toBe(true);
});
it("maps a pre/post PG-reader pair to a bounded per-second rate and nulls on NaN input", async () => {
type PgCounters = Array<{ datname: string; xactCommit: number; xactRollback: number }>;
function makeSampler(reader: () => Promise<PgCounters | null>) {
return createMetricsSampler({
runtime: inertRuntime(),
domain: { pgStatsReader: reader },
});
}
// Pre/post reader pair (per-database rows); the sampler derives a per-second rate from its
// own clock (elapsed may be ~0 on a tight loop, so the rate is finite and
// non-negative rather than a pinned value — the exact delta math is owned
// by RUFU-081's domain-sampler unit suite).
const reader = vi
.fn()
.mockResolvedValueOnce([{ datname: "fusion", xactCommit: 100, xactRollback: 0 }])
.mockResolvedValueOnce([{ datname: "fusion", xactCommit: 150, xactRollback: 0 }])
.mockResolvedValueOnce([{ datname: "fusion", xactCommit: 220, xactRollback: 0 }]);
const sampler = makeSampler(reader);
await sampler.domain.samplePgRate(); // first sample (no prior delta) -> 0
await sampler.domain.samplePgRate(); // delta over the real elapsed window
await sampler.domain.samplePgRate(); // a third delta over the next window
const rate = sampleValueOf(requireFamily(assertExpositionText(sampler.render()), "fusion_domain_postgres_queries_per_second"));
// A positive final delta yields a finite, non-negative per-second rate.
expect(Number.isFinite(rate)).toBe(true);
expect(rate).toBeGreaterThanOrEqual(0);
// NaN/unavailable input -> the field stays a finite number (0 / last-known),
// never a throw and never a NaN emitted.
const unavailable = makeSampler(async () => null);
await unavailable.domain.samplePgRate();
const nullRate = sampleValueOf(requireFamily(assertExpositionText(unavailable.render()), "fusion_domain_postgres_queries_per_second"));
expect(Number.isFinite(nullRate)).toBe(true);
});
});
describe("spawn-count hook on the real child_process module", () => {
it("increments on a real child spawn, proves the child ran, and restores the wrapper in finally", async () => {
const cp = createRequire(import.meta.url)("node:child_process") as {
spawn: (...args: unknown[]) => unknown;
};
const sampler: RuntimeSampler = createRuntimeSampler(); // defaults to the live module
const installed = sampler.installSpawnHook();
expect(installed).toBe(true);
const totalBefore = sampler.spawnCounts.total;
try {
// Spawn a REAL child through the patched module; delegate-to-original
// must let it run to completion.
const exit = await new Promise<number | null>((resolve, reject) => {
const child = cp.spawn(
process.execPath,
["-e", 'process.stdout.write("child-ran");process.exit(0)'],
{ stdio: "ignore" },
) as { on: (ev: string, cb: (code: number | null) => void) => void };
child.on("exit", (code) => resolve(code));
child.on("error", reject);
});
expect(exit).toBe(0);
// The wrap incremented the counter exactly once for this spawn.
expect(sampler.spawnCounts.total).toBeGreaterThan(totalBefore);
expect(sampler.spawnCounts.byKind.spawn).toBeGreaterThan(0);
// The sampler state reflects the observed spawn as a counter family.
const families = sampler.buildSnapshot(Date.now());
const spawnFamily = families.find((f) => f.name === "fusion_system_child_process_spawn_total");
expect(spawnFamily).toBeDefined();
expect(spawnFamily!.samples[0].value).toBeGreaterThan(0);
} finally {
// Restore the ORIGINAL functions exactly, even on assertion failure, so
// sibling tests' spawning is never poisoned.
sampler.removeSpawnHook();
}
expect(sampler.spawnHookInstalled).toBe(false);
// After removal, a direct spawn is no longer counted.
const totalAfter = sampler.spawnCounts.total;
await new Promise<void>((resolve, reject) => {
const child = cp.spawn(process.execPath, ["-e", "0"], { stdio: "ignore" }) as {
on: (ev: string, cb: (code: number | null) => void) => void;
error?: (e: Error) => void;
};
child.on("exit", () => resolve());
child.on("error", reject);
});
expect(sampler.spawnCounts.total).toBe(totalAfter);
});
it("renders the spawn counter through the orchestrator, visible to the parser", () => {
const cp = createRequire(import.meta.url)("node:child_process") as {
spawn: (...args: unknown[]) => unknown;
};
const orchestrator = createMetricsSampler();
orchestrator.runtime.installSpawnHook();
try {
cp.spawn(process.execPath, ["-e", "0"], { stdio: "ignore" });
const body = orchestrator.render();
const family = requireFamily(assertExpositionText(body), "fusion_system_child_process_spawn_total");
expect(family.samples[0].value).toBeGreaterThanOrEqual(1);
} finally {
orchestrator.runtime.removeSpawnHook();
}
});
});

View File

@@ -0,0 +1,205 @@
// @vitest-environment node
import { describe, it, expect } from "vitest";
import {
serializeMetrics,
sanitizeMetricName,
sanitizeLabelName,
escapeLabelValue,
type MetricFamily,
type MetricsSnapshot,
} from "../prometheus-text.js";
/**
* RUFU-081 Prometheus text serializer tests.
*
* Covers the pure serializer contract: scalar gauge, labeled multi-line
* family, non-finite coercion, label-value escaping, invalid-name sanitization,
* and stable HELP/TYPE + line ordering. Output must parse as a valid
* Prometheus text body (starts with `# HELP`+`# TYPE` lines, then value lines).
*/
function sampleFamilies(families: MetricFamily[]): MetricsSnapshot {
return { families };
}
function parseExpectValueLines(body: string): string[] {
const lines = body.split("\n").filter((l) => l.length > 0);
// Metadata lines start with `# `; a valid text body has no other comments.
for (const line of lines) {
expect(line.startsWith("# HELP ") || line.startsWith("# TYPE ") || !line.startsWith("#")).toBe(true);
}
const typeCount = lines.filter((l) => l.startsWith("# TYPE ")).length;
const helpCount = lines.filter((l) => l.startsWith("# HELP ")).length;
expect(typeCount).toBeGreaterThan(0);
expect(helpCount).toBe(typeCount);
return lines;
}
describe("serializeMetrics", () => {
it("emits a scalar gauge with HELP/TYPE and a stable body", () => {
const body = serializeMetrics(
sampleFamilies([
{
name: "fusion_system_cpu_seconds_total",
help: "Total user CPU seconds consumed by the server process",
type: "counter",
samples: [{ value: 12.5 }],
},
]),
);
expect(body).toBe(
["# HELP fusion_system_cpu_seconds_total Total user CPU seconds consumed by the server process", "# TYPE fusion_system_cpu_seconds_total counter", "fusion_system_cpu_seconds_total 12.5", ""].join("\n"),
);
const lines = parseExpectValueLines(body);
expect(lines).toContain("# TYPE fusion_system_cpu_seconds_total counter");
});
it("renders a labeled multi-line family with one line per label set", () => {
const body = serializeMetrics(
sampleFamilies([
{
name: "fusion_domain_tasks_total",
type: "gauge",
labels: ["column"],
samples: [
{ labelValues: ["todo"], value: 3 },
{ labelValues: ["in-progress"], value: 1 },
{ labelValues: ["in-review"], value: 0 },
],
},
]),
);
expect(body).toContain('fusion_domain_tasks_total{column="todo"} 3');
expect(body).toContain('fusion_domain_tasks_total{column="in-progress"} 1');
expect(body).toContain('fusion_domain_tasks_total{column="in-review"} 0');
expect(parseExpectValueLines(body)).toContain("# TYPE fusion_domain_tasks_total gauge");
});
it("sorts labeled samples deterministically regardless of input order", () => {
const makeBody = () =>
serializeMetrics(
sampleFamilies([
{
name: "fusion_test_labeled",
type: "gauge",
labels: ["col", "project"],
samples: [
{ labelValues: ["b", "z"], value: 1 },
{ labelValues: ["a", "a"], value: 2 },
{ labelValues: ["a", "b"], value: 3 },
],
},
]),
);
expect(makeBody()).toBe(makeBody());
const lines = makeBody().split("\n").filter((l) => l.startsWith("fusion_test_labeled"));
expect(lines).toEqual([
'fusion_test_labeled{col="a",project="a"} 2',
'fusion_test_labeled{col="a",project="b"} 3',
'fusion_test_labeled{col="b",project="z"} 1',
]);
});
it("coerces non-finite and non-numeric values to 0 instead of crashing", () => {
const body = serializeMetrics(
sampleFamilies([
{
name: "fusion_system_last_request_age_ms",
type: "gauge",
samples: [{ value: Number.NaN }],
},
{
name: "fusion_system_cpu_seconds_total",
type: "counter",
samples: [{ value: Number.POSITIVE_INFINITY }],
},
]),
);
expect(body).toContain("fusion_system_last_request_age_ms 0");
expect(body).toContain("fusion_system_cpu_seconds_total 0");
// Both families emit, no throw, and a single bad value does not abort others.
expect(parseExpectValueLines(body)).toHaveLength(6);
});
it("escapes special characters in label values", () => {
const body = serializeMetrics(
sampleFamilies([
{
name: "fusion_domain_agent_state",
type: "gauge",
labels: ["state"],
samples: [{ labelValues: ['with"quote\\and\nnewline'], value: 1 }],
},
]),
);
expect(body).toContain('fusion_domain_agent_state{state="with\\"quote\\\\and\\nnewline"} 1');
});
it("sanitizes invalid metric and label names instead of emitting broken lines", () => {
const body = serializeMetrics(
sampleFamilies([
{
name: "fusion_domain bad name",
type: "gauge",
labels: ["weird label!"],
samples: [{ labelValues: ["x"], value: 5 }],
},
]),
);
// The sanitized metric name is valid Prometheus.
expect(body).toContain("# TYPE fusion_domain_bad_name gauge");
expect(body).toContain('fusion_domain_bad_name{weird_label_="x"} 5');
const lines = parseExpectValueLines(body);
expect(lines.some((l) => l.startsWith("fusion_domain_bad_name"))).toBe(true);
});
it("handles an empty snapshot with a trailing newline only when non-empty", () => {
expect(serializeMetrics({ families: [] })).toBe("");
});
it("omits label rendering for a scalar family and emits no braces", () => {
const body = serializeMetrics(
sampleFamilies([
{ name: "fusion_system_rss_bytes", type: "gauge", samples: [{ value: 1024 }] },
]),
);
expect(body).toContain("fusion_system_rss_bytes 1024");
expect(body).not.toContain("{");
});
it("escapes newlines inside HELP text", () => {
const body = serializeMetrics(
sampleFamilies([
{ name: "fusion_test_help", type: "gauge", help: "line one\nline two", samples: [{ value: 1 }] },
]),
);
const helpLine = body.split("\n")[0];
expect(helpLine).toBe("# HELP fusion_test_help line one\\nline two");
});
});
describe("pure helpers", () => {
it("sanitizeMetricName strips forbidden characters and guards a leading digit", () => {
expect(sanitizeMetricName("fusion_system.rss")).toBe("fusion_system_rss");
expect(sanitizeMetricName("9metric")).toBe("_9metric");
expect(sanitizeMetricName("valid_name:ok")).toBe("valid_name:ok");
});
it("sanitizeLabelName strips colons and guards a leading digit", () => {
expect(sanitizeLabelName("col:umn")).toBe("col_umn");
expect(sanitizeLabelName("9label")).toBe("_9label");
expect(sanitizeLabelName("column")).toBe("column");
});
it("escapeLabelValue handles backslash, quote, and newline", () => {
expect(escapeLabelValue('a"b\\c\nd')).toBe('a\\"b\\\\c\\nd');
expect(escapeLabelValue("plain")).toBe("plain");
});
});

View File

@@ -0,0 +1,348 @@
// @vitest-environment node
import { describe, it, expect, vi, beforeEach, afterEach } from "vitest";
import express, { type Express } from "express";
import { execFileAsync } from "../../exec-file.js";
import { superviseSpawn } from "@fusion/core";
import {
createRuntimeSampler,
createRequestLatencyMiddleware,
defaultPsProbe,
} from "../runtime-sampler.js";
import type { PsProbe, PsProbeResult } from "../runtime-sampler.js";
import { get } from "../../test-request.js";
/**
* RUFU-081 runtime sampler tests.
*
* Covers the injectable surface of `createRuntimeSampler`:
* - the latency recorder measuring REAL request-pipeline durations on a live
* Express app, plus the idle freeze-indicator growing over time;
* - the bounded percentile + bucket math over the recent-duration ring;
* - the spawn hook counting spawn/fork/execFile/exec (via an injected fake
* module), being idempotent, and restoring the originals exactly;
* - the hook keeping `superviseSpawn` + `execFileAsync` working against the
* REAL `node:child_process` module after install;
* - process CPU/memory and the git-subprocess gauge via injected probes.
*/
const NODE = process.execPath;
function buildApp(sampler: ReturnType<typeof createRuntimeSampler>): Express {
const app = express();
app.use(createRequestLatencyMiddleware(sampler.latency));
app.get("/api/health", (_req, res) => {
res.json({ database: "ok" });
});
app.get("/{*splat}", (_req, res) => {
res.status(404).end();
});
return app;
}
describe("request latency recorder", () => {
let sampler: ReturnType<typeof createRuntimeSampler>;
let app: Express;
beforeEach(() => {
sampler = createRuntimeSampler();
app = buildApp(sampler);
});
it("records real request-pipeline durations on the live serving path", async () => {
const healthRes = await get(app, "/api/health");
expect(healthRes.status).toBe(200);
expect((healthRes.body as { database: string }).database).toBe("ok");
const fallbackRes = await get(app, "/some/nonexistent/path");
expect(fallbackRes.status).toBe(404);
const now = Date.now();
const families = sampler.buildSnapshot(now);
const count = families.find((f) => f.name === "fusion_system_request_count_total")!;
const latency = families.find((f) => f.name === "fusion_system_request_latency_ms")!;
const lastAge = families.find((f) => f.name === "fusion_system_last_request_age_ms")!;
expect(count.samples[0].value).toBeGreaterThanOrEqual(2);
// Every served duration is finite, non-negative.
for (const s of latency.samples) expect(Number.isFinite(s.value)).toBe(true);
// The last request completed "just now", so the freeze indicator is tiny.
expect(lastAge.samples[0].value).toBeLessThanOrEqual(60_000);
});
it("grows the last-request-age freeze indicator over idle time", () => {
sampler.recordRequest(5); // last served at Date.now()
const now = Date.now();
const families = sampler.buildSnapshot(now + 5000);
const lastAge = families.find((f) => f.name === "fusion_system_last_request_age_ms")!;
expect(lastAge.samples[0].value).toBe(5000);
});
it("never lets the freeze indicator go negative before any request", () => {
const now = Date.now();
const families = sampler.buildSnapshot(now + 5000);
const lastAge = families.find((f) => f.name === "fusion_system_last_request_age_ms")!;
expect(lastAge.samples[0].value).toBe(0);
});
it("computes bounded percentiles and bucket counts over the ring", () => {
// Push a known distribution: [10, 100, 1000].
for (const ms of [10, 100, 1000]) sampler.recordRequest(ms);
const families = sampler.buildSnapshot(0);
const latency = families.find((f) => f.name === "fusion_system_request_latency_ms")!;
const byLabel = Object.fromEntries(latency.samples.map((s) => [s.labelValues![0], s.value]));
expect(byLabel.p50).toBe(100); // 3 samples → ceil(1.5)-1 = index 1 → 100
expect(byLabel.max).toBe(1000);
const bucket = families.find((f) => f.name === "fusion_system_request_latency_bucket")!;
const le10 = bucket.samples.find((s) => s.labelValues![0] === "10")!;
expect(le10.value).toBe(1);
const le5000 = bucket.samples.find((s) => s.labelValues![0] === "5000")!;
expect(le5000.value).toBe(3);
});
});
describe("spawn hook (injected fake module)", () => {
let spawnMod: {
spawn: (...args: unknown[]) => unknown;
fork: (...args: unknown[]) => unknown;
execFile: (...args: unknown[]) => unknown;
exec: (...args: unknown[]) => unknown;
};
let sampler: ReturnType<typeof createRuntimeSampler>;
beforeEach(() => {
spawnMod = {
spawn: vi.fn(() => "child"),
fork: vi.fn(() => "child"),
execFile: vi.fn(() => "child"),
exec: vi.fn(() => "child"),
};
sampler = createRuntimeSampler({ spawnModule: spawnMod as never });
});
it("counts spawn/fork/execFile/exec then restores the originals", () => {
expect(sampler.installSpawnHook()).toBe(true);
// Call the PATCHED module functions (spawnMod.* is now wrapped). Each
// wrapped fn delegates to the original fake via .apply and increments.
spawnMod.spawn("a");
spawnMod.fork("a");
spawnMod.execFile("a");
spawnMod.exec("a");
expect(sampler.spawnCounts.total).toBe(4);
expect(sampler.spawnCounts.byKind.spawn).toBe(1);
expect(sampler.spawnCounts.byKind.fork).toBe(1);
expect(sampler.spawnCounts.byKind.execFile).toBe(1);
expect(sampler.spawnCounts.byKind.exec).toBe(1);
// Remove restores the originals (spawnMod.* is no longer wrapped).
sampler.removeSpawnHook();
const totalAfter = sampler.spawnCounts.total;
spawnMod.spawn("b");
expect(sampler.spawnCounts.total).toBe(totalAfter);
});
it("is idempotent — a second install does not double-wrap", () => {
expect(sampler.installSpawnHook()).toBe(true);
expect(sampler.installSpawnHook()).toBe(false);
spawnMod.spawn("a");
// Only one wrap: a single spawn increments by exactly 1.
expect(sampler.spawnCounts.byKind.spawn).toBe(1);
expect(sampler.spawnCounts.total).toBe(1);
});
it("renders spawn totals as counters with per-kind labels", () => {
sampler.installSpawnHook();
spawnMod.execFile("a");
const families = sampler.buildSnapshot(Date.now());
const total = families.find((f) => f.name === "fusion_system_child_process_spawn_total")!;
expect(total.type).toBe("counter");
expect(total.samples[0].value).toBe(1);
const byKind = families.find((f) => f.name === "fusion_system_child_process_spawn_total_by_kind")!;
expect(byKind.labels).toEqual(["kind"]);
expect(byKind.samples.find((s) => s.labelValues![0] === "execFile")!.value).toBe(1);
});
});
describe("spawn hook on the real child_process module", () => {
let sampler: ReturnType<typeof createRuntimeSampler>;
afterEach(() => {
sampler.removeSpawnHook();
});
it("keeps superviseSpawn and execFileAsync working after install", async () => {
sampler = createRuntimeSampler(); // defaults to the live node:child_process module
expect(sampler.installSpawnHook()).toBe(true);
const supervised = superviseSpawn(NODE, ["-e", "0"], { stdio: "ignore" });
const exit = await supervised.waitExit();
expect([0, null]).toContain(exit.code);
const { stdout } = await execFileAsync(NODE, ["-e", 'process.stdout.write("ok")']);
expect(stdout).toContain("ok");
});
});
describe("process + git gauges (injected probes)", () => {
it("renders process cpu/memory gauges from the injected process surface", async () => {
const sampler = createRuntimeSampler({
processRef: {
pid: 1234,
cpuUsage: () => ({ user: 2_000_000, system: 1_000_000 }), // 2s + 1s
memoryUsage: () => ({ rss: 1000, heapTotal: 500, heapUsed: 200 }),
},
psProbe: async () => ({ ok: false, childCommands: [], reason: "non-posix" }),
});
await sampler.sampleProcessAndGit();
const families = sampler.buildSnapshot(Date.now());
// Guard for label families that legitimately have zero samples (e.g. the
// per-kind spawn family with no spawns yet).
const byName = Object.fromEntries(families.map((f) => [f.name, f.samples[0]?.value ?? 0]));
expect(byName.fusion_system_cpu_user_seconds_total).toBe(2);
expect(byName.fusion_system_cpu_system_seconds_total).toBe(1);
expect(byName.fusion_system_process_rss_bytes).toBe(1000);
expect(byName.fusion_system_process_heap_used_bytes).toBe(200);
expect(byName.fusion_system_process_heap_total_bytes).toBe(500);
});
it("counts git child processes from a successful single-level probe", async () => {
const sampler = createRuntimeSampler({
psProbe: async () => ({ ok: true, childCommands: ["git", "git", "node"] }),
});
await sampler.sampleProcessAndGit();
const families = sampler.buildSnapshot(Date.now());
const git = families.find((f) => f.name === "fusion_system_git_child_processes")!;
expect(git.samples[0].value).toBe(2);
});
it("degrades to 0 on a failing or non-posix probe without throwing", async () => {
const failing: PsProbe = async () => ({ ok: false, childCommands: [], reason: "probe-error" });
const sampler = createRuntimeSampler({ psProbe: failing });
await sampler.sampleProcessAndGit();
const families = sampler.buildSnapshot(Date.now());
const git = families.find((f) => f.name === "fusion_system_git_child_processes")!;
expect(git.samples[0].value).toBe(0);
expect(Number.isFinite(git.samples[0].value)).toBe(true);
});
it("defaultPsProbe returns a settling result on any platform", async () => {
// Regardless of platform, the probe always settles with a finite outcome.
const result = await defaultPsProbe(process.pid);
expect(typeof result.ok).toBe("boolean");
expect(Array.isArray(result.childCommands)).toBe(true);
});
});
describe("overlap guard (RUFU-081 Greptile P1 #2)", () => {
/*
FNXC:MetricsSampler 2026-08-18-11:53 (RUFU-081 CodeRabbit Major, RUFU-106 review fix):
An async `sampleProcessAndGit` that outlasts its arm's interval must never overlap the next tick
under the same guard key. These fake-timer tests hold the `psProbe` promise pending while a
second interval fires and assert the probe is invoked once (the tick was skipped). The `process`
and `git` arms SHARE one guard key (both call sampleProcessAndGit), so a coinciding tick on the
other arm is skipped too instead of double-probing (CodeRabbit Major review fix 2026-08-18-11:53).
*/
it("skips a process tick still in flight — the ps probe is invoked at most once per completed window", async () => {
vi.useFakeTimers();
let resolveProbe: (r: PsProbeResult) => void = () => {};
const pending = new Promise<PsProbeResult>((res) => {
resolveProbe = res;
});
const psProbe = vi.fn(() => pending);
const sampler = createRuntimeSampler({ tick: { processMs: 5000, gitMs: 1_000_000 }, psProbe });
sampler.start();
// First window fires -> probe invoked, sample stays pending (in flight).
await vi.advanceTimersByTimeAsync(5000);
expect(psProbe).toHaveBeenCalledTimes(1);
// Second window fires while tick 1 is still awaiting -> SKIPPED (probe NOT re-invoked).
await vi.advanceTimersByTimeAsync(5000);
expect(psProbe).toHaveBeenCalledTimes(1);
// Resolve the in-flight sample; the `finally` clears the guard.
resolveProbe({ ok: false, childCommands: [], reason: "probe-error" });
await vi.advanceTimersByTimeAsync(0);
// Guard clear -> process arm fires again, exactly once per window.
await vi.advanceTimersByTimeAsync(5000);
expect(psProbe).toHaveBeenCalledTimes(2);
sampler.stopTimers();
});
it("skips a git tick still in flight — the git arm never overlaps its own in-flight sample", async () => {
vi.useFakeTimers();
let resolveProbe: (r: PsProbeResult) => void = () => {};
const pending = new Promise<PsProbeResult>((res) => {
resolveProbe = res;
});
const psProbe = vi.fn(() => pending);
// processMs far out, so only the git (15s) arm fires in this window.
const sampler = createRuntimeSampler({ tick: { processMs: 1_000_000, gitMs: 15_000 }, psProbe });
sampler.start();
await vi.advanceTimersByTimeAsync(15_000);
expect(psProbe).toHaveBeenCalledTimes(1);
// Second git window fires while the first is still awaiting -> SKIPPED.
await vi.advanceTimersByTimeAsync(15_000);
expect(psProbe).toHaveBeenCalledTimes(1);
resolveProbe({ ok: false, childCommands: [], reason: "probe-error" });
await vi.advanceTimersByTimeAsync(0);
await vi.advanceTimersByTimeAsync(15_000);
expect(psProbe).toHaveBeenCalledTimes(2);
sampler.stopTimers();
});
/*
* FNXC:MetricsSampler 2026-08-18-11:53 (RUFU-081 CodeRabbit Major, RUFU-106 review fix):
* With the DEFAULT 5000/15000 ms cadence the process and git arms coincide at t=15000; the
* shared guard must skip the coinciding git tick instead of launching a second concurrent
* `ps` probe. This test keeps both arms at their defaults (no pushed-out interval).
*/
it("skips the coinciding git tick under the shared guard at default cadence — one probe per completed window", async () => {
vi.useFakeTimers();
let resolveProbe: (r: PsProbeResult) => void = () => {};
const pending = new Promise<PsProbeResult>((res) => {
resolveProbe = res;
});
const psProbe = vi.fn(() => pending);
// Default cadence: process 5000 ms, git 15000 ms — the arms coincide at t=15000.
const sampler = createRuntimeSampler({ psProbe });
sampler.start();
// t=5000: the process arm fires -> probe #1 (still pending).
await vi.advanceTimersByTimeAsync(5000);
expect(psProbe).toHaveBeenCalledTimes(1);
// t=10000: the process tick is skipped (same guard key, in flight).
await vi.advanceTimersByTimeAsync(5000);
expect(psProbe).toHaveBeenCalledTimes(1);
// t=15000: the git tick coincides while the process sample is still pending ->
// the SHARED guard skips it; no second probe.
await vi.advanceTimersByTimeAsync(5000);
expect(psProbe).toHaveBeenCalledTimes(1);
// Resolve the in-flight sample; the shared guard clears in `finally`.
resolveProbe({ ok: false, childCommands: [], reason: "probe-error" });
await vi.advanceTimersByTimeAsync(0);
// t=20000: the process arm fires again -> probe #2.
await vi.advanceTimersByTimeAsync(5000);
expect(psProbe).toHaveBeenCalledTimes(2);
sampler.stopTimers();
});
afterEach(() => {
vi.useRealTimers();
});
});

View File

@@ -0,0 +1,465 @@
/**
* RUFU-081 domain sampler: PostgreSQL query-rate gauge + Fusion-domain gauges
* (project / running-agent / board-column counts) for the dashboard `/metrics`
* endpoint.
*
* The PG rate closes the "poll-storm" diagnosis gap: `pg_stat_database`
* xact_commit/xact_rollback delta, normalized to a per-second rate on the
* tick. The baseline tracks counters PER DATABASE so a stats reset in one
* database is detected even when the cross-database sum stays positive,
* and a failed-probe gap invalidates the baseline so a reset that lands
* inside the gap can never produce a cross-epoch (fabricated) rate. The Fusion-domain gauges answer the "how busy is the board" questions
* (active/paused project split, running agents, tasks per column) from stores
* the dashboard has ALREADY opened — never by opening a store, starting an
* engine, or starting a watcher just to answer a scrape.
*
* Non-blocking by design, matching `runtime-sampler.ts`:
* - every read happens on a pre-read tick; `buildSnapshot` performs zero
* awaited I/O and only renders pre-read numeric state;
* - the PG read is best-effort and degrades (keeps the last-known rate, or
* `0` on the first invalid sample) when the layer is absent, the DB is
* privilege-fenced (embedded PG from agent sessions), or a pool error
* occurs — it never throws and never hammers the DB;
* - duplicate/undefined project ids are deduped so they never produce
* duplicate or malformed metric lines.
*
* All seams are injectable so tests exercise the sampler without booting
* stores or a database (mirroring `createRuntimeSampler(init)`).
*/
import { drizzleSql } from "@fusion/core";
import type { MetricFamily } from "./prometheus-text.js";
const sql = drizzleSql;
import { listRegisteredProjectStores, countRunningAgentsInStore } from "../project-store-resolver.js";
/** Per-database cumulative counters for one `pg_stat_database` probe row. */
export interface PgDbStats {
/** `pg_stat_database.datname` — per-DB identity for baseline tracking. */
datname: string;
/** Cumulative transactions committed since the last stats reset (this database). */
xactCommit: number;
/** Cumulative transactions rolled back since the last stats reset (this database). */
xactRollback: number;
}
/** One probe sample: the per-database cumulative counters (one entry per `pg_stat_database` row). */
export type PgStats = PgDbStats[];
/** Reads the per-database cumulative PG counters, or resolves `null` when no layer exists. */
export type PgStatsReader = () => Promise<PgStats | null>;
/**
* A slim task view: only the column is needed to count cards per column.
* Kept narrow so tests can hand over `{ column }` objects without a real store.
*/
export interface SlimTaskLike {
column: string;
}
/** Registry that yields the already-open project stores (injectable). */
export type RegisteredStoreRegistry = () => Array<{ projectId: string; store: unknown }>;
/** The pre-read domain snapshot a `/metrics` render reflects. */
export interface DomainSamplerState {
/** Per-second transaction delta, or 0/last-known when unavailable. */
pgQueriesPerSecond: number;
/** Registered-project split by running-agent activity. */
projectCounts: { total: number; active: number; idle: number };
/** projectId -> number of running agents in that project's open store. */
runningAgentsByProject: Record<string, number>;
/** columnId -> number of tasks in that column across all registered projects. */
columnCounts: Record<string, number>;
}
/** Constructor options (all injectable). */
export interface DomainSamplerInit {
/** Source of already-open project stores (defaults to `listRegisteredProjectStores`). */
registeredStores?: RegisteredStoreRegistry;
/** Per-store running-agent counter (defaults to `countRunningAgentsInStore`). */
countAgentsInStore?: (store: unknown) => Promise<number>;
/** Per-store slim task listing (defaults to `store.listTasks({ slim: true })`). */
listTasksInStore?: (store: unknown) => Promise<SlimTaskLike[]>;
/** PG cumulative-counter reader (defaults to the store-layer `pg_stat_database` probe). */
pgStatsReader?: PgStatsReader;
/** Tick cadences in ms (defaults: PG 5s / domain 5s). */
tick?: { pgMs?: number; domainMs?: number };
/** A fake-timer-friendly `setInterval`/`clearInterval` surface. */
timers?: {
setInterval: (fn: () => void, ms: number) => { unref?: () => void };
clearInterval: (t: { unref?: () => void }) => void;
};
}
/** The domain sampler's public handle. */
export interface DomainSampler {
readonly state: DomainSamplerState;
readonly started: boolean;
/** Read the PG rate + domain gauges into the pre-read snapshot now. */
samplePgRate(): Promise<void>;
/** Read the domain gauges into the pre-read snapshot now. */
sampleDomain(): Promise<void>;
/** Start interval timers (unref'd so they never keep the process alive). */
start(): void;
/** Clear interval timers and fence out any in-flight sample. */
stopTimers(): void;
/** Assemble the domain metric families for a scrape (synchronous, O(metric count)). */
buildSnapshot(nowMs?: number): MetricFamily[];
}
/** Per-database cumulative counter baseline values (one entry per `pg_stat_database.datname`). */
type PgBaseline = Map<string, { xactCommit: number; xactRollback: number }>;
/** Default timers from the global scope (fake-timer injectable). */
function defaultTimers(): DomainSamplerInit["timers"] {
return {
setInterval: (fn, ms) => setInterval(fn, ms) as unknown as { unref?: () => void },
clearInterval: (t) => clearInterval(t as unknown as ReturnType<typeof setInterval>),
};
}
/** Default registry: the dashboard's already-open project-store cache. */
function defaultRegistry(): RegisteredStoreRegistry {
return () => listRegisteredProjectStores() as Array<{ projectId: string; store: unknown }>;
}
/** Default per-store running-agent counter. */
async function defaultCountAgents(store: unknown): Promise<number> {
return countRunningAgentsInStore(store as Parameters<typeof countRunningAgentsInStore>[0]);
}
/** Default slim task listing through the store's public `listTasks({ slim: true })`. */
async function defaultListTasks(store: unknown): Promise<SlimTaskLike[]> {
const anyStore = store as { listTasks?: (opts: { slim: true }) => Promise<SlimTaskLike[]> };
if (typeof anyStore.listTasks !== "function") return [];
return anyStore.listTasks({ slim: true });
}
/** Default PG cumulative-counter reader. Reads `pg_stat_database` (one row per database) on the
* first registered store that exposes a live async layer. Best-effort: a missing layer,
* privilege-fenced PG, or pool error resolves `null` so the sampler keeps the last-known rate
* instead of throwing.
*/
export function defaultPgStatsReader(): PgStatsReader {
return async (): Promise<PgStats | null> => {
const stores = listRegisteredProjectStores();
for (const { store } of stores) {
const layer = (store as { getAsyncLayer?: () => { db: { execute: (q: unknown) => Promise<unknown[]> } } | null }).getAsyncLayer?.();
if (!layer) continue;
try {
const rows = (await layer.db.execute(
sql.raw(`
SELECT datname, xact_commit, xact_rollback FROM pg_stat_database
`),
)) as Array<{ datname: string; xact_commit: number; xact_rollback: number }>;
if (!rows || rows.length === 0) return null;
return rows.map((row) => ({
datname: String(row.datname ?? ""),
xactCommit: Number(row.xact_commit) || 0,
xactRollback: Number(row.xact_rollback) || 0,
}));
} catch {
// Privilege-fenced / transient — degrade to null, never throw.
return null;
}
}
return null;
};
}
/** Create a domain sampler. Timers are NOT started until {@link DomainSampler.start}. */
export function createDomainSampler(init: DomainSamplerInit = {}): DomainSampler {
const registry = init.registeredStores ?? defaultRegistry();
const countAgents = init.countAgentsInStore ?? defaultCountAgents;
const listTasks = init.listTasksInStore ?? defaultListTasks;
const pgReader = init.pgStatsReader ?? defaultPgStatsReader();
const tick = { pgMs: init.tick?.pgMs ?? 5000, domainMs: init.tick?.domainMs ?? 5000 };
const timers = init.timers ?? defaultTimers();
const state: DomainSamplerState = {
pgQueriesPerSecond: 0,
projectCounts: { total: 0, active: 0, idle: 0 },
runningAgentsByProject: {},
columnCounts: {},
};
let previousPg: PgBaseline | undefined;
let previousPgAtMs: number | null = null;
/*
* FNXC:MetricsSampler 2026-08-18-04:20 (RUFU-081 Greptile P1, RUFU-106 review fix):
* A failed probe marks the retained baseline STALE: a stats reset landing inside the failed
* gap is invisible to the next successful probe (the counter may already have grown past the
* retained total, so the cross-epoch delta would read positive), so the first success after a
* failed gap re-baselines and keeps the last-known rate instead of emitting a fabricated rate.
*/
let pgBaselineStale = false;
let pgEverBaselined = false;
/*
* FNXC:MetricsSampler 2026-08-18-04:20 (RUFU-081 Greptile P1, RUFU-106 review fix):
* Restart fencing: stopTimers() bumps the generation, and a sample that started before the
* stop (its pre-close read still awaiting) must not write its stale result after the sampler
* restarts. The in-flight guard below is factory-scoped so it SURVIVES restarts — a pre-close
* sample still running at restart keeps blocking new ticks until it resolves (and is then
* fenced out of the write).
*/
let sampleGeneration = 0;
const inFlight = new Set<string>();
let started = false;
const timersMap = new Map<string, { unref?: () => void }>();
async function samplePgRate(): Promise<void> {
const gen = sampleGeneration;
let stats: PgStats | null = null;
try {
stats = await pgReader();
} catch {
stats = null;
}
// Restarted mid-sample: discard this read — its write would be stale (Greptile P1 review fix).
if (gen !== sampleGeneration) return;
const now = Date.now();
if (!stats || stats.length === 0) {
/*
* FNXC:MetricsSampler 2026-08-16-23:35 (RUFU-081 Greptile P1 #1, RUFU-106):
* A transient reader failure (null or thrown) must NOT reset the PG baseline. Resetting
* `previousPg = undefined` here made the next good sample look like a FIRST sample and
* report rate 0, even though real queries kept flowing. On failure we keep the last-known
* `previousPg` and `state.pgQueriesPerSecond`; only a successful counter read
* establishes/advances the baseline (a first-ever failure still leaves the rate at 0).
*/
pgBaselineStale = true;
return;
}
const current: PgBaseline = new Map();
for (const row of stats) {
if (typeof row?.datname !== "string" || row.datname.length === 0) continue;
current.set(row.datname, {
xactCommit: Number(row.xactCommit) || 0,
xactRollback: Number(row.xactRollback) || 0,
});
}
if (current.size === 0) {
pgBaselineStale = true;
return;
}
if (!previousPg || !previousPgAtMs || pgBaselineStale) {
previousPg = current;
previousPgAtMs = now;
pgBaselineStale = false;
if (!pgEverBaselined) {
state.pgQueriesPerSecond = 0;
pgEverBaselined = true;
}
// A re-baseline after a stale gap keeps the last-known rate (see pgBaselineStale comment).
return;
}
const elapsedMs = now - previousPgAtMs;
let delta = 0;
let reset = false;
for (const [datname, cur] of current) {
const prev = previousPg.get(datname);
if (!prev) continue; // database appeared mid-window: join the baseline, no delta yet
const d = cur.xactCommit + cur.xactRollback - (prev.xactCommit + prev.xactRollback);
if (d < 0) {
/*
* FNXC:MetricsSampler 2026-08-18-04:20 (RUFU-081 Greptile P1, RUFU-106 review fix):
* A PER-DB counter went backward: a stats reset in this database. The cross-database sum
* can stay positive when other databases grew in the same window, so an aggregate-only
* reset check would accept the cross-epoch delta — the per-DB check is the only detector.
*/
reset = true;
break;
}
delta += d;
}
previousPg = current;
previousPgAtMs = now;
pgBaselineStale = false;
if (reset || elapsedMs <= 0) {
// Clock skew or a stats reset (pg_stat_reset) — keep the last-known rate.
return;
}
state.pgQueriesPerSecond = (delta / elapsedMs) * 1000;
}
async function sampleDomain(): Promise<void> {
const gen = sampleGeneration;
// Registry entries are keyed by projectId; dedupe so a duplicate/undefined
// project id can never produce duplicate or malformed metric lines.
const entries = new Map<string, { store: unknown }>();
for (const entry of registry()) {
const projectId = entry?.projectId;
if (typeof projectId !== "string" || projectId.length === 0) continue;
// First registration of a project id is canonical; later duplicates are
// dropped so a duplicate id can never double-count a column or an agent.
if (!entries.has(projectId)) entries.set(projectId, { store: entry.store });
}
// Use the deduped keys so both per-project and aggregate counts agree.
const projectIds = [...entries.keys()];
const runningByProject: Record<string, number> = {};
const columnCounts: Record<string, number> = {};
await Promise.all(
projectIds.map(async (projectId) => {
const { store } = entries.get(projectId)!;
let running = 0;
try {
running = await countAgents(store);
} catch {
running = 0;
}
runningByProject[projectId] = running;
let tasks: SlimTaskLike[] = [];
try {
tasks = await listTasks(store);
} catch {
tasks = [];
}
for (const task of tasks) {
// Each slim task row contributes one card to its column. Column ids
// are strings; guard against undefined/malformed rows.
const column = task?.column;
if (typeof column !== "string" || column.length === 0) continue;
columnCounts[column] = (columnCounts[column] ?? 0) + 1;
}
}),
);
// Restarted mid-sample: discard this read — its write would be stale (Greptile P1 review fix).
if (gen !== sampleGeneration) return;
const total = projectIds.length;
const active = projectIds.filter((id) => (runningByProject[id] ?? 0) > 0).length;
state.runningAgentsByProject = runningByProject;
state.columnCounts = columnCounts;
state.projectCounts = { total, active, idle: total - active };
}
function start(): void {
if (started) return;
started = true;
/*
* FNXC:MetricsSampler 2026-08-17-01:01 (RUFU-081 Greptile P1 #2, RUFU-106):
* In-flight flags are FACTORY-SCOPED (not created here) so the guard survives a
* stopTimers()+start() restart: a sample still running from before the restart keeps
* blocking ticks of the same arm until it resolves, where it is fenced out of its write
* by the generation check (Greptile P1 review fix 2026-08-18).
*/
const arm = (key: string, intervalMs: number, run: () => Promise<void>): void => {
const timer = timers!.setInterval(() => {
/*
* FNXC:MetricsSampler 2026-08-17-01:01 (RUFU-081 Greptile P1 #2, RUFU-106):
* An async sample that outlasts its interval must never overlap the next tick. If this
* sampler's previous run is still awaiting, skip the tick entirely; otherwise set the flag,
* run the sample, and clear it in `finally` so the next interval arm is armed again.
*/
if (inFlight.has(key)) return;
inFlight.add(key);
void run()
.catch(() => undefined)
.finally(() => inFlight.delete(key));
}, intervalMs);
timer.unref?.();
timersMap.set(key, timer);
};
arm("pg", tick.pgMs, samplePgRate);
arm("domain", tick.domainMs, sampleDomain);
}
function stopTimers(): void {
started = false;
// Fence out any in-flight sample: it started before this stop, so its write (if it
// resolves after a restart) must be discarded (Greptile P1 review fix 2026-08-18).
sampleGeneration += 1;
/*
* FNXC:MetricsSampler 2026-08-18-11:53 (RUFU-081 Greptile P1, RUFU-106 review fix):
* The retained PG baseline is STALE on stop too: if PostgreSQL statistics reset during
* the stop gap, counters regrow past the retained totals and the first post-restart
* success would emit a cross-epoch delta as a positive rate. Marking the baseline stale
* makes the first post-restart success re-baseline and keep the last-known rate.
*/
pgBaselineStale = true;
for (const key of [...timersMap.keys()]) {
const timer = timersMap.get(key);
if (timer) {
try {
timers!.clearInterval(timer);
} catch {
/* ignore */
}
timersMap.delete(key);
}
}
}
function buildSnapshot(_nowMs?: number): MetricFamily[] {
const families: MetricFamily[] = [];
// ── PostgreSQL query rate ──────────────────────────────────────────────
families.push({
name: "fusion_domain_postgres_queries_per_second",
help: "PostgreSQL transaction rate (xact_commit + xact_rollback delta per second); degrades to last-known when unreadable",
type: "gauge",
samples: [{ value: state.pgQueriesPerSecond }],
});
// ── Project + running-agent counts ─────────────────────────────────────
families.push({
name: "fusion_domain_projects_total",
help: "Registered open projects (already-open stores only)",
type: "gauge",
samples: [{ value: state.projectCounts.total }],
});
families.push({
name: "fusion_domain_projects_active",
help: "Registered projects currently running at least one agent",
type: "gauge",
samples: [{ value: state.projectCounts.active }],
});
families.push({
name: "fusion_domain_projects_idle",
help: "Registered projects running zero agents (idle/unpaused park)",
type: "gauge",
samples: [{ value: state.projectCounts.idle }],
});
const agentProjectSamples = Object.entries(state.runningAgentsByProject).map(([projectId, count]) => ({
labelValues: [projectId],
value: count,
}));
families.push({
name: "fusion_domain_project_running_agents",
help: "Running agents per registered project",
type: "gauge",
labels: ["project"],
samples: agentProjectSamples,
});
// ── Board task counts per column ───────────────────────────────────────
const columnSamples = Object.entries(state.columnCounts).map(([column, count]) => ({
labelValues: [column],
value: count,
}));
families.push({
name: "fusion_domain_board_tasks",
help: "Tasks per board column across all registered projects",
type: "gauge",
labels: ["column"],
samples: columnSamples,
});
return families;
}
return {
state,
get started() {
return started;
},
samplePgRate,
sampleDomain,
start,
stopTimers,
buildSnapshot,
};
}

View File

@@ -0,0 +1,39 @@
/**
* RUFU-081 combined `/metrics` observability surface.
*
* Single import point for the Prometheus-text serializer and the runtime /
* domain samplers that feed the dashboard `/metrics` endpoint. The orchestrator
* ({@link createMetricsSampler}) is what `server.ts` mounts; the lower-level
* modules are re-exported for tests and future samplers.
*/
export { serializeMetrics } from "./prometheus-text.js";
export type { MetricSample, MetricFamily, MetricsSnapshot } from "./prometheus-text.js";
export { createMetricsSampler } from "./sampler.js";
export type { MetricsSampler, MetricsSamplerInit } from "./sampler.js";
export {
createRuntimeSampler,
createRequestLatencyMiddleware,
recordRequest,
defaultPsProbe,
REQUEST_LATENCY_RING_CAP,
DEFAULT_LATENCY_BUCKETS_MS,
} from "./runtime-sampler.js";
export type {
RuntimeSampler,
RuntimeSamplerInit,
LatencyRecorderState,
SpawnCounts,
ProcessLike,
PsProbe,
} from "./runtime-sampler.js";
export { createDomainSampler, defaultPgStatsReader } from "./domain-sampler.js";
export type {
DomainSampler,
DomainSamplerInit,
DomainSamplerState,
PgStats,
PgStatsReader,
} from "./domain-sampler.js";

View File

@@ -0,0 +1,165 @@
/**
* Pure Prometheus text exposition serializer for the dashboard `/metrics`
* endpoint (RUFU-081).
*
* This module has ZERO side effects and never reads the clock, the network, or
* the environment. Callers pass an explicit snapshot (a list of metric
* families) and get back a deterministic, scrapable Prometheus text body.
*
* Why direct text serialization instead of `prom-client` or a conversion from
* the OTLP wire shape:
* - A `/metrics` scrape is a plain `curl`/Prometheus scrape; the endpoint IS
* the surface. No new third-party metric library is added.
* - The OTLP mapping (`packages/core/src/process/otel-metrics.ts`) produces
* the collector wire shape ({@link OtlpExportPayload}); this module reuses
* only its gauge/counter *semantics* (point-in-time gauges vs monotonic
* counters), not its wire envelope.
*
* Invariants enforced here:
* - Output is deterministic (stable family + line order) so a diff of two
* scrapes only shows real changes.
* - Bad values never crash a scrape: non-finite / non-numeric values are
* coerced to `0` (documented choice) so one NaN cannot take down the whole
* body.
* - Label values are escaped per the exposition format (`\\`, `\"`, `\n`, `\`).
* - Invalid metric/label names are sanitized to the permitted character set
* rather than rejected, so a ragged runtime value never sabotages the body.
*/
/** A single value or label-keyed value line for a metric family. */
export interface MetricSample {
/** Serialize values stably by sorting on this key first when present. */
labelValues?: string[];
/** The numeric value; non-finite/non-numeric is coerced to 0. */
value: number;
}
/** A Prometheus metric family (one HELP/TYPE pair plus sample lines). */
export interface MetricFamily {
/** Prometheus metric name. Sanitized on serialize if invalid. */
name: string;
/** Human-readable HELP text (never reproduced in the run-audit). */
help?: string;
/** `gauge` (point-in-time) or `counter` (monotonic) semantics. */
type: "gauge" | "counter";
/**
* Label names shared by every sample line in the family. Provide
* `labels` AND per-sample `labelValues` (same cardinality) for a labeled
* family; omit both for a scalar family.
*/
labels?: string[];
/** One line per sample. For a labeled family each entry contributes one label set. */
samples: MetricSample[];
}
/** A full scrape snapshot assembled from pre-read gauge state. */
export interface MetricsSnapshot {
families: MetricFamily[];
}
/** Permitted Prometheus metric-name characters: `[a-zA-Z_:][a-zA-Z0-9_:]*`. */
const METRIC_NAME_RE = /^[a-zA-Z_:][a-zA-Z0-9_:]*$/;
/** Permitted Prometheus label-name characters: `[a-zA-Z_][a-zA-Z0-9_]*`. */
const LABEL_NAME_RE = /^[a-zA-Z_][a-zA-Z0-9_]*$/;
/** A trailing run of forbidden characters, used to sanitize metric names. */
const INVALID_METRIC_CHARS_RE = /[^a-zA-Z0-9_:]/g;
const INVALID_LABEL_CHARS_RE = /[^a-zA-Z0-9_]/g;
/**
* Coerce a value to a finite number, defaulting to `0`. NaN, Infinity,
* undefined, null, and strings that don't parse numerically all become `0` so
* a single bad sample can never abort the whole exposition body.
*/
function coerceValue(value: number): number {
const numeric = typeof value === "number" ? value : Number(value);
return Number.isFinite(numeric) ? numeric : 0;
}
/**
* Sanitize a metric name to the permitted character set. Colons are reserved
* for recording rules / client libraries and are preserved here because the
* sampler chooses valid names at the call site; the sanitizer only strips
* characters the exposition format forbids so a ragged runtime string can
* never produce an unparseable body.
*/
export function sanitizeMetricName(name: string): string {
const cleaned = String(name).replace(INVALID_METRIC_CHARS_RE, "_");
return METRIC_NAME_RE.test(cleaned) ? cleaned : `_${cleaned}`;
}
/**
* Sanitize a label name. Same contract as {@link sanitizeMetricName} but for
* the narrower label-name character set (no colon).
*/
export function sanitizeLabelName(name: string): string {
const cleaned = String(name).replace(INVALID_LABEL_CHARS_RE, "_");
return LABEL_NAME_RE.test(cleaned) ? cleaned : `_${cleaned}`;
}
/**
* Escape a label value per the Prometheus text exposition format: `\` ->
* `\\`, `"` -> `\"`, and newline -> `\n`. All other bytes pass through.
*/
export function escapeLabelValue(value: string): string {
return String(value)
.replace(/\\/g, "\\\\")
.replace(/"/g, '\\"')
.replace(/\n/g, "\\n");
}
/**
* Render a single label set `{name="value",other="value"}` (the leading brace
* inclusive). Returns an empty string for a scalar (no-labels) family.
*/
function renderLabelSet(labels: string[], labelValues: string[]): string {
if (labels.length === 0) return "";
const parts = labels.map((rawName, index) => {
const name = sanitizeLabelName(rawName);
const rawValue = labelValues[index];
const value = escapeLabelValue(rawValue === undefined ? "" : rawValue);
return `${name}="${value}"`;
});
return `{${parts.join(",")}}`;
}
/**
* Serialize a snapshot into a Prometheus text body.
*
* Ordering is deterministic: families are emitted in the order given (the
* sampler owns the "meaningful order" contract — runtime first, domain
* second), and within a labeled family samples are sorted by their joined
* label values so scraping with `?sort=` stability is not required to diff
* scrapes. Each family contributes exactly one `# HELP` and one `# TYPE` line
* followed by its value lines.
*/
export function serializeMetrics(snapshot: MetricsSnapshot): string {
const lines: string[] = [];
for (const family of snapshot.families) {
const name = sanitizeMetricName(family.name);
const help = family.help ?? `${name} measurement`;
const type = family.type;
// HELP/TYPE lines. HELP text is newline-escaped so a description with a
// line break cannot inject a spurious line into the body.
lines.push(`# HELP ${name} ${escapeLabelValue(help)}`);
lines.push(`# TYPE ${name} ${type}`);
if (family.labels && family.labels.length > 0) {
const labeled = family.samples
.map((sample) => ({ sample, key: (sample.labelValues ?? []).join("\u0000") }))
.sort((a, b) => (a.key < b.key ? -1 : a.key > b.key ? 1 : 0));
for (const { sample } of labeled) {
const value = coerceValue(sample.value);
lines.push(`${name}${renderLabelSet(family.labels, sample.labelValues ?? [])} ${value}`);
}
} else {
for (const sample of family.samples) {
const value = coerceValue(sample.value);
lines.push(`${name} ${value}`);
}
}
}
return `${lines.join("\n")}${lines.length > 0 ? "\n" : ""}`;
}

View File

@@ -0,0 +1,594 @@
/**
* RUFU-081 runtime sampler: request-path latency recorder, process CPU/memory
* gauges, child-process spawn counter, and a bounded git-subprocess gauge.
*
* This module feeds the Prometheus-text `/metrics` endpoint
* (`prometheus-text.ts`). It owns the in-process measurement state that a
* `/metrics` scrape must render synchronously from a pre-read snapshot — a
* scrape or sampler tick must NEVER starve the event loop, so every sampler is
* cheap and non-blocking:
* - request latency is recorded per served request in O(1) and stored in a
* bounded ring;
* - process CPU/memory come from the synchronous `process.cpuUsage()` /
* `process.memoryUsage()` calls;
* - the spawn counter is a monkey-patch that delegates to the bound original
* via `.call(this, ...)` so child spawning is never broken and nested usage
* through `superviseSpawn` / `runCommandAsync` keeps working;
* - the git-subprocess gauge is a single-level `ps --ppid <pid>` (POSIX-only,
* best-effort, at most every ~15s, NEVER recursive).
*
* The whole module is intentionally framework-light: every dependency is
* injectable (process, child_process, ps probe, timers) so samplers and the
* spawn hook are unit-testable without spawning real children or booting
* Express. The latency recorder exposes an Express-compatible
* `(req, res, next)` middleware (see {@link createRequestLatencyMiddleware}).
*/
/*
FNXC:MetricsEndpoint 2026-08-13-17:38:
RUFU-081: serve a Prometheus-text `GET /metrics` on the dashboard exposing the
five system/runtime/Fusion-domain measurements a 2026-08-13 CPU/UI-freeze
diagnosis collected by hand (event-loop health latency, native spawn cadence,
PG query rate, git subprocess count, engine CPU/RSS). This is the RUNTIME
sampler half: request-path latency (the ONLY direct freeze indicator), process
CPU/memory gauges, the child_process spawn-count hook, and the bounded
git-subprocess probe.
Non-negotiable constraints honored here:
- A /metrics scrape or sampler tick must NEVER starve the event loop: the
handler renders synchronously from pre-read state, and every sampler is
cheap/best-effort.
- The event-loop/health-latency metric must reflect the LIVE serving path
(the actual HTTP handler cost), not a synthetic probe. That is why the
latency recorder is an Express middleware mounted before route handlers and
records `finish` (real pipeline cost) on the response.
- The spawn hook patches the live CommonJS exports object of
node:child_process (not a statically-destructured function ref) so native
`spawn@:-1` callers observable via `require`/`import * as cp` are counted;
it delegates via .apply and NEVER breaks child spawning.
- The git gauge is a single-level `ps --ppid <pid>` (POSIX-only, best-effort,
~15s, never recursive) that degrades to 0 on failure.
- Metric values are numeric gauges; nothing here writes prose into the
run-audit (FN-7158/FN-7528). No GitHub push; this lands via local main /
operator only.
*/
import { exec } from "node:child_process";
import { createRequire } from "node:module";
import type { MetricFamily } from "./prometheus-text.js";
/**
* The live CommonJS exports object of `node:child_process`. Monkey-patching
* this OBJECT (instead of a statically-destructured function reference) makes
* the spawn-count hook visible to `require("node:child_process")` and
* `import * as cp` callers — the call-time object-access path the 2026-08-13
* diagnosis saw as native `spawn@:-1` frames with no JS parent. Statically
* destructured callers (e.g. `import { spawn }`) capture the original at module
* load and are not re-routed — an accepted, inherent limit. The hook NEVER
* breaks child spawning: the wrapper delegates to the original via `.apply`.
*/
const childProcessRequire = createRequire(import.meta.url);
/** Cap on the number of recent request durations retained for histogram math. */
export const REQUEST_LATENCY_RING_CAP = 256;
/** Default histogram bucket edges in milliseconds (Prometheus histogram). */
export const DEFAULT_LATENCY_BUCKETS_MS = [5, 10, 25, 50, 100, 250, 500, 1000, 2500, 5000];
/** Spawn-hook gauge label for the operation kind. */
const SPAWN_KIND_LABEL = "kind";
/** A process surface the sampler reads gauges from (injectable for tests). */
export interface ProcessLike {
pid: number;
cpuUsage: (prev?: { user: number; system: number }) => { user: number; system: number };
memoryUsage: () => {
rss: number;
heapTotal: number;
heapUsed: number;
};
}
/**
* A hookable process-command surface. `spawn`, `fork`, `execFile`, and `exec`
* are patched by the spawn counter and restored on stop. The wrapper delegates
* to the ORIGINAL via `.call(this, ...)` so receiver-bound callers and nested
* usage through `superviseSpawn` / `runCommandAsync` keep working.
*/
export interface SpawnableChildProcessModule {
spawn: (...args: unknown[]) => unknown;
fork: (...args: unknown[]) => unknown;
execFile: (...args: unknown[]) => unknown;
exec: (...args: unknown[]) => unknown;
}
/** Numeric shape the spawn counter derives from the hook. */
export interface SpawnCounts {
/** Total spawn/fork/execFile/exec invocations since the hook was installed. */
total: number;
/** Per-kind cumulative counts keyed by `"spawn" | "fork" | "execFile" | "exec"`. */
byKind: Record<string, number>;
}
/** Mutable state the latency recorder updates per served request. */
export interface LatencyRecorderState {
/** Ring of the most recent served-request durations (ms). */
recent: number[];
/** Max duration (ms) seen since the last reset/start. */
maxMs: number;
/** Monotonic cumulative request count. */
requestCount: number;
/** Epoch ms of the last served request, or -1 if none served yet. */
lastServedAtMs: number;
}
/** A stubbed `ps --ppid` result: exit code and parsed child rows. */
export interface PsProbeResult {
ok: boolean;
/** Child process command names from a single-level `ps --ppid <pid>`. */
childCommands: string[];
/** When `ok` is false, set to a short reason like "ENOENT" | "non-posix". */
reason?: string;
}
/**
* A callable `ps` probe. Kept injectable so tests can substitute a fake
* without spawning a real process. The production default runs a bounded,
* single-level `ps -o comm= --ppid <pid>` scan.
*/
export type PsProbe = (pid: number) => Promise<PsProbeResult>;
/** A fake-timer-friendly interval surface. */
export interface TimerLike {
unref?: () => void;
}
/** The runtime sampler handles samplers, hooks, and snapshot rendering. */
export interface RuntimeSampler {
/** Latency recorder state (bounded ring + last-served timestamp). */
readonly latency: LatencyRecorderState;
/** Spawn counter state (cumulative + per-kind). */
readonly spawnCounts: SpawnCounts;
/** True while the spawn hook is installed. */
readonly spawnHookInstalled: boolean;
/** True while any interval timer is running. */
readonly started: boolean;
/** Install the spawn-count hook (idempotent). Returns true if newly installed. */
installSpawnHook(): boolean;
/** Remove the spawn-count hook, restoring the original functions exactly. */
removeSpawnHook(): void;
/** Record one served request duration (ms) into the bounded ring. */
recordRequest(ms: number): void;
/** Sample process CPU/memory + git-subprocess gauges now. */
sampleProcessAndGit(): Promise<void>;
/** Start interval timers (unref'd so they never keep the process alive). */
start(): void;
/** Clear interval timers; does NOT remove the spawn hook. */
stopTimers(): void;
/** Assemble the runtime metric families for a scrape (synchronous, O(N)). */
buildSnapshot(nowMs?: number): MetricFamily[];
}
/** Constructor options (all injectable for tests). */
export interface RuntimeSamplerInit {
/** Process API surface (defaults to the global `process`). */
processRef?: ProcessLike;
/** child_process surface to patch (defaults to the live Node module). */
spawnModule?: SpawnableChildProcessModule;
/** Injectable `ps --ppid` probe (defaults to the real single-level scan). */
psProbe?: PsProbe;
/** Latency histogram bucket edges in ms (defaults to {@link DEFAULT_LATENCY_BUCKETS_MS}). */
latencyBucketsMs?: number[];
/** Ring cap for recent request durations (defaults to {@link REQUEST_LATENCY_RING_CAP}). */
ringCap?: number;
/** Tick cadences in ms (defaults: latency 5s / process 5s / git 15s). */
tick?: { latencyMs?: number; processMs?: number; gitMs?: number };
/** A fake-timer-friendly `setInterval`/`clearInterval` surface. */
timers?: {
setInterval: (fn: () => void, ms: number) => TimerLike;
clearInterval: (t: TimerLike) => void;
};
}
/**
* Build an Express-style `(req, res, next)` request-latency recorder head.
* Must be mounted on the app BEFORE route handlers so it times the LIVE serving
* path (including `GET /api/health`), never a synthetic probe. The recorder
* attaches a `finish` listener on the response (when it has one) so it measures
* the full request pipeline cost, and always calls `next()`.
*/
export function createRequestLatencyMiddleware(state: LatencyRecorderState) {
return (
// req/res are typed loosely: the recorder only needs the response `once`.
_req: unknown,
res: unknown,
next?: () => void,
): void => {
const startedAt = Date.now();
if (isResponseLike(res)) {
res.once("finish", () => {
recordRequest(state, Date.now() - startedAt);
});
} else {
// Non-response contexts (unit tests): record synchronously.
recordRequest(state, Date.now() - startedAt);
}
if (typeof next === "function") {
next();
}
};
}
/** true when the object behaves like an HTTP response (`once` method). */
function isResponseLike(res: unknown): res is { once: (event: string, cb: () => void) => unknown } {
return (
typeof res === "object" && res !== null && typeof (res as { once?: unknown }).once === "function"
);
}
/** O(1) bound-ring insert for one served-request duration. */
export function recordRequest(
state: LatencyRecorderState,
ms: number,
ringCap = REQUEST_LATENCY_RING_CAP,
): void {
const capped = Number.isFinite(ms) && ms >= 0 ? ms : 0;
state.recent.push(capped);
if (state.recent.length > ringCap) {
state.recent.shift();
}
if (capped > state.maxMs) state.maxMs = capped;
state.requestCount += 1;
state.lastServedAtMs = Date.now();
}
/** Compute a percentile over the recent durations ring (0 when empty). */
function percentile(recent: number[], pct: number): number {
if (recent.length === 0) return 0;
const sorted = [...recent].sort((a, b) => a - b);
const idx = Math.min(sorted.length - 1, Math.max(0, Math.ceil((pct / 100) * sorted.length) - 1));
return sorted[idx] ?? 0;
}
/** Build the bounded histogram bucket counts over the recent ring. */
function buildBuckets(recent: number[], edges: number[]): Record<string, number> {
const counts: Record<string, number> = {};
for (const edge of edges) {
counts[String(edge)] = recent.filter((d) => d <= edge).length;
}
return counts;
}
/** The production single-level `ps -o comm= --ppid <pid>` probe (POSIX-only). */
export function defaultPsProbe(pid: number): Promise<PsProbeResult> {
const platform = typeof process !== "undefined" ? process.platform : "posix";
if (platform === "win32") {
return Promise.resolve({ ok: false, childCommands: [], reason: "non-posix" });
}
return new Promise((resolve) => {
exec(`ps -o comm= --ppid ${Number(pid)}`, { timeout: 2000, maxBuffer: 1024 * 1024 }, (error, stdout) => {
if (error) {
resolve({ ok: false, childCommands: [], reason: "probe-error" });
return;
}
const commands = stdout
.split("\n")
.map((line) => line.trim())
.filter((line) => line.length > 0 && line !== "COMMAND");
resolve({ ok: true, childCommands: commands });
});
});
}
/** Default process surface adapter for the production `process`. */
function defaultProcessLike(): ProcessLike {
return {
pid: process.pid,
cpuUsage: () => process.cpuUsage(),
memoryUsage: () => process.memoryUsage(),
};
}
/** Default timers from the global scope (fake-timer injectable). */
function defaultTimers(): {
setInterval: (fn: () => void, ms: number) => TimerLike;
clearInterval: (t: TimerLike) => void;
} {
return {
setInterval: (fn, ms) => setInterval(fn, ms) as unknown as TimerLike,
clearInterval: (t) => clearInterval(t as unknown as ReturnType<typeof setInterval>),
};
}
/**
* Create a runtime sampler.
*
* The spawn hook is NOT installed until {@link RuntimeSampler.installSpawnHook}
* is called; the interval timers are NOT started until
* {@link RuntimeSampler.start} is called. `start()` uses `unref()`'d timers so
* a running sampler never keeps the process alive.
*/
export function createRuntimeSampler(init: RuntimeSamplerInit = {}): RuntimeSampler {
const proc = init.processRef ?? defaultProcessLike();
const spawnMod = init.spawnModule ??
(childProcessRequire("node:child_process") as SpawnableChildProcessModule);
const psProbe = init.psProbe ?? defaultPsProbe;
const buckets = init.latencyBucketsMs ?? DEFAULT_LATENCY_BUCKETS_MS;
const ringCap = init.ringCap ?? REQUEST_LATENCY_RING_CAP;
const tick = {
latencyMs: init.tick?.latencyMs ?? 5000,
processMs: init.tick?.processMs ?? 5000,
gitMs: init.tick?.gitMs ?? 15_000,
};
const timers = init.timers ?? defaultTimers();
const latencyState: LatencyRecorderState = {
recent: [],
maxMs: 0,
requestCount: 0,
lastServedAtMs: -1,
};
const spawnCounts: SpawnCounts = { total: 0, byKind: {} };
// Hold the ORIGINAL functions so stop() can restore them exactly.
const originalSpawnFns: Partial<SpawnableChildProcessModule> = {};
let spawnHookInstalled = false;
let started = false;
// Last sampled process gauges + git subprocess count.
let lastCpu: { user: number; system: number } = { user: 0, system: 0 };
let lastMem: { rss: number; heapTotal: number; heapUsed: number } = { rss: 0, heapTotal: 0, heapUsed: 0 };
let lastGitCount = 0;
// ── Interval timer slots ────────────────────────────────────────────────
const timersMap = new Map<string, TimerLike>();
// ── Spawn counter hook ──────────────────────────────────────────────────
const SPAWN_KINDS: Array<keyof SpawnableChildProcessModule> = ["spawn", "fork", "execFile", "exec"];
function installSpawnHook(): boolean {
if (spawnHookInstalled) return false;
for (const kind of SPAWN_KINDS) {
const original = spawnMod[kind];
if (typeof original !== "function") continue;
originalSpawnFns[kind] = original;
const wrapped = function (this: unknown, ...args: unknown[]) {
spawnCounts.total += 1;
spawnCounts.byKind[kind] = (spawnCounts.byKind[kind] ?? 0) + 1;
// Delegate to the ORIGINAL via .apply(this, ...) so receiver-bound
// callers and nested usage through superviseSpawn/runCommandAsync keep
// working.
return (original as (...a: unknown[]) => unknown).apply(this, args);
} as typeof original;
(spawnMod as unknown as Record<string, unknown>)[kind] = wrapped;
}
spawnHookInstalled = true;
return true;
}
function removeSpawnHook(): void {
if (!spawnHookInstalled) return;
for (const kind of SPAWN_KINDS) {
const original = originalSpawnFns[kind];
if (original !== undefined) {
(spawnMod as unknown as Record<string, unknown>)[kind] = original;
}
delete originalSpawnFns[kind];
}
spawnHookInstalled = false;
}
// ── Process + git sample ────────────────────────────────────────────────
async function sampleProcessAndGit(): Promise<void> {
try {
lastCpu = proc.cpuUsage();
} catch {
lastCpu = { user: 0, system: 0 };
}
try {
lastMem = proc.memoryUsage();
} catch {
lastMem = { rss: 0, heapTotal: 0, heapUsed: 0 };
}
try {
const result = await psProbe(proc.pid);
if (result.ok && result.childCommands.some((c) => c.toLowerCase().startsWith("git"))) {
lastGitCount = result.childCommands.filter((c) => c.toLowerCase().startsWith("git")).length;
} else {
// Degrade to 0 rather than throwing; a ps failure is not a scrape error.
lastGitCount = 0;
}
} catch {
lastGitCount = 0;
}
}
// ── Interval timers ─────────────────────────────────────────────────────
function start(): void {
if (started) return;
started = true;
// Per-sampler in-flight flags so a tick that fires while the previous sample is still
// awaiting is SKIPPED: samplers never run concurrently and a slow sample never queues
// (RUFU-081 Greptile P1 #2, RUFU-106). The `process` and `git` arms both invoke
// sampleProcessAndGit, so they SHARE one guard key ("process-git"): with the default
// 5000/15000 ms cadence the arms coincide every 15 seconds and independent keys would
// launch two concurrent `ps` probes on the coinciding tick (CodeRabbit Major review fix
// 2026-08-18-11:53).
const inFlight = new Set<string>();
const arm = (key: string, guardKey: string, intervalMs: number, run: () => Promise<void>): void => {
const timer = timers.setInterval(() => {
/*
* FNXC:MetricsSampler 2026-08-17-01:01 (RUFU-081 Greptile P1 #2, RUFU-106):
* An async sample that outlasts its interval must never overlap the next tick under the
* same guard key. If a run is still awaiting, skip the tick; otherwise set the flag,
* run the sample, and clear it in `finally` so the next interval fires again.
*/
if (inFlight.has(guardKey)) return;
inFlight.add(guardKey);
void run()
.catch(() => {
/* best-effort */
})
.finally(() => inFlight.delete(guardKey));
}, intervalMs);
// Best-effort unref; fake timers may not expose it, but default timers
// are unref'd so a running sampler never keeps the process alive.
timer.unref?.();
timersMap.set(key, timer);
};
arm("latency", "latency", tick.latencyMs, () => {
// The latency "sampler" tick is a no-op marker: the useful measurements
// already live in the ring from actual served requests.
return Promise.resolve();
});
arm("process", "process-git", tick.processMs, sampleProcessAndGit);
arm("git", "process-git", tick.gitMs, sampleProcessAndGit);
}
function stopTimers(): void {
started = false;
for (const key of [...timersMap.keys()]) {
const timer = timersMap.get(key);
if (timer) {
try {
timers.clearInterval(timer);
} catch {
/* ignore */
}
timersMap.delete(key);
}
}
}
// ── Snapshot ────────────────────────────────────────────────────────────
function buildSnapshot(nowMs?: number): MetricFamily[] {
const now = nowMs ?? Date.now();
const families: MetricFamily[] = [];
// ── Request latency / event-loop health ────────────────────────────────
const lastAgeMs = latencyState.lastServedAtMs >= 0 ? Math.max(0, now - latencyState.lastServedAtMs) : 0;
families.push({
name: "fusion_system_request_count_total",
help: "Total HTTP requests served through the latency recorder",
type: "counter",
samples: [{ value: latencyState.requestCount }],
});
families.push({
name: "fusion_system_request_latency_ms",
help: "Served request latency in milliseconds over the recent ring",
type: "gauge",
labels: ["quantile"],
samples: [
{ labelValues: ["p50"], value: percentile(latencyState.recent, 50) },
{ labelValues: ["p95"], value: percentile(latencyState.recent, 95) },
{ labelValues: ["max"], value: latencyState.maxMs },
],
});
const bucketCounts = buildBuckets(latencyState.recent, buckets);
families.push({
name: "fusion_system_request_latency_bucket",
help: "Cumulative count of served requests at or below the bucket edge (ms)",
type: "gauge",
labels: ["le"],
samples: Object.entries(bucketCounts).map(([edge, count]) => ({
labelValues: [edge],
value: count,
})),
});
families.push({
name: "fusion_system_last_request_age_ms",
help: "Milliseconds since the last served request; grows during event-loop starvation (freeze indicator)",
type: "gauge",
samples: [{ value: lastAgeMs }],
});
// ── Process CPU / memory ───────────────────────────────────────────────
// cpuUsage() returns microseconds; expose seconds for a human-friendly gauge.
families.push({
name: "fusion_system_cpu_user_seconds_total",
help: "User CPU time consumed by the server process (cumulative seconds)",
type: "counter",
samples: [{ value: lastCpu.user / 1_000_000 }],
});
families.push({
name: "fusion_system_cpu_system_seconds_total",
help: "System CPU time consumed by the server process (cumulative seconds)",
type: "counter",
samples: [{ value: lastCpu.system / 1_000_000 }],
});
families.push({
name: "fusion_system_process_rss_bytes",
help: "Resident set size of the server process",
type: "gauge",
samples: [{ value: lastMem.rss }],
});
families.push({
name: "fusion_system_process_heap_used_bytes",
help: "Heap used by the server process",
type: "gauge",
samples: [{ value: lastMem.heapUsed }],
});
families.push({
name: "fusion_system_process_heap_total_bytes",
help: "Total heap allocated to the server process",
type: "gauge",
samples: [{ value: lastMem.heapTotal }],
});
// ── Child-process spawn counters ───────────────────────────────────────
families.push({
name: "fusion_system_child_process_spawn_total",
help: "Cumulative child_process spawn/fork/execFile/exec invocations since hook install",
type: "counter",
samples: [{ value: spawnCounts.total }],
});
const kindSamples = Object.entries(spawnCounts.byKind).map(([kind, count]) => ({
labelValues: [kind],
value: count,
}));
// Emit the per-kind family only when at least one kind has been observed, so
// an empty spawn hook never produces an empty-samples family (malformed
// output). The scalar total above is always present (0 before any spawn).
if (kindSamples.length > 0) {
families.push({
name: "fusion_system_child_process_spawn_total_by_kind",
help: "Cumulative child_process spawn/fork/execFile/exec invocations by kind",
type: "counter",
labels: [SPAWN_KIND_LABEL],
samples: kindSamples,
});
}
// ── Git subprocess gauge ───────────────────────────────────────────────
families.push({
name: "fusion_system_git_child_processes",
help: "Live git child processes of the server process (single-level ps --ppid, best-effort)",
type: "gauge",
samples: [{ value: lastGitCount }],
});
return families;
}
return {
latency: latencyState,
spawnCounts,
get spawnHookInstalled() {
return spawnHookInstalled;
},
get started() {
return started;
},
installSpawnHook,
removeSpawnHook,
recordRequest: (ms: number) => recordRequest(latencyState, ms, ringCap),
sampleProcessAndGit,
start,
stopTimers,
buildSnapshot,
};
}

View File

@@ -0,0 +1,104 @@
/**
* RUFU-081 sampler orchestrator for the dashboard `/metrics` endpoint.
*
* Composes the runtime sampler (request-latency recorder, process CPU/memory,
* spawn-count hook, git-subprocess gauge) and the domain sampler (PostgreSQL
* query rate + project/agent/board-column gauges) behind one handle:
*
* - `start()` — installs the spawn-count hook once and starts both samplers'
* unref'd tick timers (runtime ~5s, process ~5s, git ~15s,
* PG ~5s, domain ~5s). Idempotent: a second `start()` after the
* first is a no-op and never stacks a second wrap of
* `child_process`.
* - `stop()` — clears all timers and removes the spawn hook, exactly
* restoring the original `child_process` functions.
* - `render()` — assembles the Prometheus text body SYNCHRONOUSLY from the
* pre-read gauge snapshot (zero awaited I/O in the render
* path, so a scrape can never starve the event loop).
* - `middleware()` — the latency-recorder head for `app.use(...)`, mounted
* before route handlers so it times the LIVE HTTP serving
* path (including `GET /api/health`).
*
* No GitHub push, no publish/release/tag commands are ever run by this module.
* Metric values are numeric gauges only; nothing here writes metric content or
* numeric snapshots into the run-audit (FN-7158/FN-7528).
*/
import type { MetricFamily } from "./prometheus-text.js";
import { serializeMetrics } from "./prometheus-text.js";
import {
createRuntimeSampler,
createRequestLatencyMiddleware,
type RuntimeSampler,
type RuntimeSamplerInit,
} from "./runtime-sampler.js";
import { createDomainSampler, type DomainSampler, type DomainSamplerInit } from "./domain-sampler.js";
/** The orchestrator's public handle. */
export interface MetricsSampler {
readonly runtime: RuntimeSampler;
readonly domain: DomainSampler;
/** True while the sampler timers + spawn hook are active. */
readonly started: boolean;
/**
* The Express latency-recorder head. Mount with `app.use(...)` before route
* handlers; it records every served request into the shared ring.
*/
middleware(): (req: unknown, res: unknown, next?: () => void) => void;
/** Install the spawn hook once + start all tick timers (idempotent). */
start(): void;
/** Clear all timers and remove the spawn hook (idempotent). */
stop(): void;
/** Render the Prometheus text body synchronously from pre-read state. */
render(nowMs?: number): string;
}
/** Constructor options; both sampler configs are fully injectable for tests. */
export interface MetricsSamplerInit {
runtime?: RuntimeSamplerInit;
domain?: DomainSamplerInit;
}
/** Create an orchestrator. No side effects until {@link MetricsSampler.start}. */
export function createMetricsSampler(init: MetricsSamplerInit = {}): MetricsSampler {
const runtime = createRuntimeSampler(init.runtime);
const domain = createDomainSampler(init.domain);
let started = false;
function start(): void {
if (started) return;
started = true;
// installSpawnHook is itself idempotent; start once so a repeat start never
// stacks a second wrapper over child_process.
runtime.installSpawnHook();
runtime.start();
domain.start();
}
function stop(): void {
if (!started) return;
started = false;
runtime.stopTimers();
domain.stopTimers();
runtime.removeSpawnHook();
}
function render(nowMs?: number): string {
// Synchronous render from pre-read gauges only — no awaits here.
const now = nowMs ?? Date.now();
const families: MetricFamily[] = [...runtime.buildSnapshot(now), ...domain.buildSnapshot(now)];
return serializeMetrics({ families });
}
return {
runtime,
domain,
get started() {
return started;
},
middleware: () => createRequestLatencyMiddleware(runtime.latency),
start,
stop,
render,
};
}

View File

@@ -0,0 +1,228 @@
// @vitest-environment node
import { EventEmitter } from "node:events";
import { describe, expect, it, vi } from "vitest";
import type { Settings, TaskStore } from "@fusion/core";
import { assertExpositionText, indexFamilies, requireFamily } from "../../__tests__/prometheus-text-parse.js";
import { createServer } from "../../server.js";
import { request } from "../../test-request.js";
/*
FNXC:PrometheusAcceptance 2026-08-13-16:45:
RUFU-082 endpoint acceptance: the served /metrics body must be proven well-formed
Prometheus exposition text that covers all five measurement gaps RUFU-081
introduced (event-loop/health latency, spawn cadence, PG query rate, git gauge,
CPU/memory/RSS) and must NOT be the SPA index.html fallback that /metrics used to
serve. A scrape must stay store-free (no run-audit prose writes) and repeat scrapes
must render a fresh, bounded snapshot. Assertions bind to RUFU-081's actual metric
family names, never assumed ones.
*/
/**
* RUFU-082 endpoint acceptance suite for the `/metrics` observability route.
*
* RUFU-081 implemented and serialized this route; RUFU-082 is the acceptance
* grammar that proves the served body is REAL Prometheus exposition text — and
* that the ORIGINAL bug class is gone: `GET /metrics` used to fall through to
* the SPA shell and serve `index.html`. These tests parse the served body with
* an independent exposition-text parser (see `../../__tests__/prometheus-text-parse.ts`)
* and assert, per family, the five diagnosis gaps the RUFU-081 contract
* enumerates, plus the no-run-audit-prose and repeat-scrape-stability seams.
*
* Because the sampler is created inside `createServer` and only observable via
* the route, this suite drives the app the same way a `curl /metrics` scrape
* would — through the real request pipeline via `test-request.ts`. No
* production DB queries, no polling, no real network.
*
* Five measurement gaps asserted here:
* 1. event-loop/health latency -> fusion_system_request_latency_ms /
* fusion_system_last_request_age_ms
* 2. child-process spawn count -> fusion_system_child_process_spawn_total
* 3. PostgreSQL query rate -> fusion_domain_postgres_queries_per_second
* 4. git subprocess gauge -> fusion_system_git_child_processes
* 5. engine CPU / memory / RSS -> fusion_system_cpu_{user,system}_seconds_total /
* fusion_system_process_{rss,heap_*}_bytes
*/
/** Minimal store double (mirrors the app-level route test fixtures). */
class MockStore extends EventEmitter {
getRootDir(): string {
return "/repo";
}
getFusionDir(): string {
return "/repo/.fusion";
}
// FNXC:PostgresCutover: server setup probes the async layer, so the route
// double exposes the production-shaped backend seam.
getAsyncLayer = vi.fn(() => ({
db: {
update: vi.fn(() => ({
set: vi.fn(() => ({
where: vi.fn(() => ({ returning: vi.fn(async () => []) })),
})),
})),
},
}));
getSettings = vi.fn(async () => this.getSettingsFast());
getSettingsFast = vi.fn(async (): Promise<Settings> => ({} as Settings));
getTaskWorkflowSelection = vi.fn(() => undefined);
getWorkflowDefinition = vi.fn(async () => undefined);
getWorkflowSettingValues = vi.fn(() => ({}));
getWorkflowSettingsProjectId = vi.fn(() => "default");
getProjectScopedPluginMcpServers = vi.fn().mockResolvedValue([]);
// The run-audit write seam. A scrape must NEVER write metric prose/content
// here (FN-7158/FN-7528); the test asserts this spy is untouched during a
// GET /metrics/N.
recordRunAuditEvent = vi.fn();
}
function createApp(store: MockStore) {
return createServer(store as unknown as TaskStore, { noAuth: true });
}
describe("GET /metrics (RUFU-082 acceptance)", () => {
it("serves parseable Prometheus exposition text, not the SPA index.html fallback", async () => {
const app = createApp(new MockStore());
const res = await request(app, "GET", "/metrics");
expect(res.status).toBe(200);
const contentType = String(res.headers["content-type"] ?? "");
expect(contentType).toContain("text/plain");
expect(contentType).toContain("version=0.0.4");
const body = String(res.body);
// The original bug class: this path used to fall through to the SPA shell.
expect(body).not.toContain("<!doctype html>");
expect(body).not.toContain("<html");
expect(body).not.toContain('id="root"');
// Independent grammar proves it is well-formed exposition text.
const parsed = assertExpositionText(body);
expect(parsed.families.length).toBeGreaterThan(0);
});
it.each([
["event-loop/health latency", "fusion_system_request_latency_ms"],
["event-loop/health freeze indicator", "fusion_system_last_request_age_ms"],
["child-process spawn count", "fusion_system_child_process_spawn_total"],
["postgresql query rate", "fusion_domain_postgres_queries_per_second"],
["git subprocess gauge", "fusion_system_git_child_processes"],
["engine CPU (user seconds)", "fusion_system_cpu_user_seconds_total"],
["engine process RSS", "fusion_system_process_rss_bytes"],
["engine heap used", "fusion_system_process_heap_used_bytes"],
])("exposes the %s measurement gap as a parseable family", async (_label, familyName) => {
const app = createApp(new MockStore());
const res = await request(app, "GET", "/metrics");
expect(res.status).toBe(200);
const parsed = assertExpositionText(String(res.body));
const family = requireFamily(parsed, familyName);
// A family must be present and its scalar (no-label) value a finite number
// in the fresh-process empty state (the serializer coerces NaN/Inf to 0).
expect(family.samples.length).toBeGreaterThan(0);
const value = family.samples[0].value;
expect(Number.isFinite(value)).toBe(true);
});
it("ties the event-loop latency family to the LIVE serving path (a real request moves the gauges)", async () => {
const app = createApp(new MockStore());
// A real request through the pipeline (the health route) must be recorded
// by the latency middleware and reflected in the next scrape.
const health = await request(app, "GET", "/api/health");
expect(health.status).toBe(200);
const res = await request(app, "GET", "/metrics");
const parsed = assertExpositionText(String(res.body));
const index = indexFamilies(parsed);
const count = requireFamily(parsed, "fusion_system_request_count_total").samples[0].value;
expect(count).toBeGreaterThanOrEqual(1);
// The last-request-age freeze indicator reflects the just-served request
// (a small, finite number) rather than the pre-request 0.
const age = requireFamily(parsed, "fusion_system_last_request_age_ms").samples[0].value;
expect(Number.isFinite(age)).toBe(true);
expect(age).toBeLessThan(5000);
// The latency quantile family exposes the labeled p50/p95/max quantiles.
const latency = index.get("fusion_system_request_latency_ms");
expect(latency).toBeDefined();
const quantiles = Object.fromEntries(
latency!.samples.map((s) => [s.labels[0]?.value, s.value]),
);
for (const q of ["p50", "p95", "max"]) {
expect(quantiles).toHaveProperty(q);
expect(Number.isFinite(quantiles[q])).toBe(true);
}
});
it("writes no run-audit row (metric prose/content) during a scrape", async () => {
const store = new MockStore();
const app = createApp(store);
// A first request warms server construction (some setup paths touch the
// store); then we care only about the scrape itself being store-free.
await request(app, "GET", "/api/health");
store.recordRunAuditEvent.mockClear();
const res = await request(app, "GET", "/metrics");
expect(res.status).toBe(200);
const body = String(res.body);
// The scrape is a pure synchronous render from pre-read gauges; it must
// never emit audit rows. This guards the FN-7158/FN-7528 "no prose in
// run-audit" invariant and that metric values are numeric-only.
expect(store.recordRunAuditEvent).not.toHaveBeenCalled();
// The body itself must be numeric gauges only — no JSON/text prose lines.
const parsed = assertExpositionText(body);
for (const sample of parsed.samples) {
expect(typeof sample.value).toBe("number");
}
});
it("a second immediate scrape is fresh/parseable with a stable, bounded family count", async () => {
const app = createApp(new MockStore());
const first = await request(app, "GET", "/metrics");
const parsedFirst = assertExpositionText(String(first.body));
const familyCountFirst = parsedFirst.families.length;
// A second immediate scrape must not pile up ever-growing per-scrape series
// (each render is a fresh snapshot from bounded pre-read state).
const second = await request(app, "GET", "/metrics");
const parsedSecond = assertExpositionText(String(second.body));
expect(parsedSecond.families.length).toBe(familyCountFirst);
// The per-family sample counts stay identical too (same set, same sizes).
const shapeFirst = parsedFirst.families.map((f) => [f.name, f.samples.length]);
const shapeSecond = parsedSecond.families.map((f) => [f.name, f.samples.length]);
expect(shapeSecond).toEqual(shapeFirst);
});
it("does not confuse an adjacent path that serves non-metric content (fallback guard)", async () => {
const app = createApp(new MockStore());
// A navigation path we did not turn into a metrics route must still be
// served by the SPA fallback (index shell) and NOT by the /metrics handler.
const spaRes = await request(app, "GET", "/some/navigation/path");
const spaBody = String(spaRes.body);
// The SPA fallback serves the shell (an HTML boot page) for navigation
// paths — in test mode that is a "temporarily unavailable" boot page rather
// than a static index.html. What matters is it is NOT Prometheus text.
expect(spaBody).not.toContain("# TYPE fusion_system_request_count_total");
expect(spaBody).not.toContain("# HELP fusion_system_request_count_total");
expect(spaBody).not.toContain("text/plain; version=0.0.4");
// The real /metrics route is unaffected and still parses.
const metricsRes = await request(app, "GET", "/metrics");
expect(metricsRes.status).toBe(200);
assertExpositionText(String(metricsRes.body));
});
it("empty-state (fresh process) renders well-formed zero/NaN-coerced scalar families", async () => {
const app = createApp(new MockStore());
const parsed = assertExpositionText(String((await request(app, "GET", "/metrics")).body));
// Fresh process: PG rate is 0 (no prior delta), git gauge 0, spawn count 0.
expect(requireFamily(parsed, "fusion_domain_postgres_queries_per_second").samples[0].value).toBe(0);
expect(requireFamily(parsed, "fusion_system_git_child_processes").samples[0].value).toBe(0);
expect(requireFamily(parsed, "fusion_system_child_process_spawn_total").samples[0].value).toBe(0);
// The labeled per-kind spawn family is absent when nothing spawned (the
// serializer omits an empty-samples labeled family), which is well-formed.
expect(indexFamilies(parsed).has("fusion_system_child_process_spawn_total_by_kind")).toBe(false);
});
});

View File

@@ -94,6 +94,7 @@ import {
} from "./reliability-metrics.js";
import { loadViewChunkManifest, type ViewChunkManifestEntry } from "./view-chunk-manifest.js";
import { maybeStartOtelExporter, type OtelExporterHandle } from "./otel-exporter.js";
import { createMetricsSampler } from "./metrics/index.js";
import { requireAsyncLayer } from "./require-async-layer.js";
import {
evaluateDashboardPostgresHealth,
@@ -978,6 +979,16 @@ export function createServer(store: TaskStore, options?: ServerOptions): ReturnT
const app = express();
app.locals.hybridExecutor = options?.hybridExecutor;
/*
FNXC:MetricsEndpoint 2026-08-13-16:15:
RUFU-081: per-server /metrics observability. The orchestrator is created with
no side effects here; the spawn-count hook + tick timers start only on listen
and stop on close (co-located with the OTLP exporter). It runs in both
headless and non-headless servers. Its latency-recorder middleware is mounted
below, before route handlers, so it times the LIVE serving path.
*/
const metricsSampler = createMetricsSampler();
const runtimeLogger = options?.runtimeLogger ?? createRuntimeLogger("server");
const mutationRateLimit = rateLimit(RATE_LIMITS.mutation);
const setupRateLimit = rateLimit(RATE_LIMITS.api);
@@ -1026,6 +1037,16 @@ export function createServer(store: TaskStore, options?: ServerOptions): ReturnT
});
});
/*
FNXC:MetricsEndpoint 2026-08-13-16:15:
RUFU-081: mount the request-latency recorder head on every request (inside
/api and the SPA shell, headless or not) so it measures the real HTTP serving
pipeline cost, including /api/health — the single best event-loop-starvation
indicator. It only attaches a `finish` listener and calls next(); it never
blocks or serializes the render path.
*/
app.use(metricsSampler.middleware());
// Daemon mode: bearer token authentication middleware
// Auth is enabled when daemon option is provided OR FUSION_DAEMON_TOKEN env var is set.
// The middleware exempts /api/health and everything outside /api/ — the SPA shell
@@ -1845,6 +1866,21 @@ export function createServer(store: TaskStore, options?: ServerOptions): ReturnT
}));
});
/*
FNXC:MetricsEndpoint 2026-08-13-16:15:
RUFU-081: the /metrics route is mounted at the app level (NOT under /api) so
it is public and scrapable like the SPA shell — daemon bearer-token auth only
protects /api/*. This is intentional: the body is pre-read numeric gauges
only (no secrets, no prose), served synchronously from the sampler snapshot
with zero awaited I/O so a scrape can never starve the event loop or itself
be subject to on-demand DB/ps work. It must be mounted before the SPA
catch-all below so it returns Prometheus text rather than index.html.
*/
app.get("/metrics", (_req, res) => {
res.type("text/plain; version=0.0.4; charset=utf-8");
res.send(metricsSampler.render());
});
app.get("/api/engine/status", (req, res) => {
const projectId = getProjectIdFromRequest(req);
res.json(buildEngineStatusPayload(projectId, options));
@@ -2372,6 +2408,22 @@ export function createServer(store: TaskStore, options?: ServerOptions): ReturnT
});
}
// RUFU-081: start the /metrics samplers (spawn-count hook + unref'd tick
// timers). Synchronous and best-effort; a failure here must never break
// server startup or the request pipeline. The guard IS the failure
// isolation the comment promises — `runtime.installSpawnHook()` patches
// node:child_process members, and an unguarded throw would propagate out
// of listen() and abort startup (CodeRabbit Minor review fix 2026-08-18-11:53,
// matching the adjacent OTLP exporter pattern).
try {
metricsSampler.start();
} catch (error) {
runtimeLogger.warn("Metrics sampler failed to start", {
message: "Metrics sampler failed to start",
...normalizeErrorForLog(error),
});
}
if (!providerHealthMonitor && (options?.engineManager || options?.engine)) {
const providerHealthLogger = runtimeLogger.child("provider-health");
providerHealthMonitor = new ProviderHealthMonitor({
@@ -2395,6 +2447,18 @@ export function createServer(store: TaskStore, options?: ServerOptions): ReturnT
aiSessionStore?.stopScheduledCleanup();
otelExporter?.stop();
otelExporter = null;
// RUFU-081: stop the /metrics samplers and remove the spawn hook so no
// timer or wrapper outlives the server on restart/test teardown. Guarded
// so a teardown throw cannot skip providerHealthMonitor?.stop() and the
// remaining close handlers (CodeRabbit Minor review fix 2026-08-18-11:53).
try {
metricsSampler.stop();
} catch (error) {
runtimeLogger.warn("Metrics sampler failed to stop", {
message: "Metrics sampler failed to stop",
...normalizeErrorForLog(error),
});
}
providerHealthMonitor?.stop();
providerHealthMonitor = null;
(apiRouter as Router & { dispose?: () => void }).dispose?.();