FN-7796: switch Grok CLI headless prompts to the reliable single-object JSON output

Narrative: the streaming-json headless contract intermittently emitted only thought events then stopReason:Cancelled with zero text, leaving Chat replies silently empty; the adapter now spawns grok with --output-format json, buffers stdout, and parses the single JSON response on process close, with streaming-json parsing kept only as a diagnostic fallback.

- Change grok CLI invocation from --output-format streaming-json to --output-format json (cli-stream.ts)
- Add GrokCliJsonResponse type ({text, stopReason, sessionId, requestId, thought}) and parseJsonOutput() to stream-parser.ts, keeping legacy NDJSON line parsing for fallback/diagnostics
- Rework runtime-adapter.ts to buffer full stdout, parse it via parsePromptOutput (JSON object first, NDJSON fallback), and surface a formatTerminalNoTextDiagnostic when a non-EndTurn stopReason yields no assistant text
- Rename first-line/inactivity timeout bookkeeping from line-based to output/chunk-based (FIRST_OUTPUT_TIMEOUT_MS, firstOutputReceived, firstStdoutChunk) since stdout is no longer consumed via readline
- Update cli-stream/runtime-adapter/stream-parser tests to cover the JSON response path and the Cancelled/no-text diagnostic
- Update docs/grok-cli-contract.md and plugin README to document the json output-format contract and diagnostics
- Add changeset fn-7796-grok-cli-reliable-headless.md (patch, fix)

Files changed:
 .changeset/fn-7796-grok-cli-reliable-headless.md   |   7 +
 docs/grok-cli-contract.md                          | 108 ++++++++-----
 plugins/fusion-plugin-grok-runtime/README.md       |  16 +-
 .../src/__tests__/cli-stream.test.ts               |   4 +-
 .../src/__tests__/runtime-adapter.test.ts          |  73 ++++++++-
 .../src/__tests__/stream-parser.test.ts            |  80 +++++----
 .../fusion-plugin-grok-runtime/src/cli-stream.ts   |  14 +-
 .../src/runtime-adapter.ts                         | 180 ++++++++++++---------
 .../src/stream-parser.ts                           |  65 ++++++--
 plugins/fusion-plugin-grok-runtime/src/types.ts    |  13 +-
 10 files changed, 373 insertions(+), 187 deletions(-)

Fusion-Task-Id: FN-7796

Fusion-Task-Lineage: c920fcf0-98f8-42ec-867a-7f76c0aca1b7

Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
This commit is contained in:
gsxdsm
2026-07-10 12:57:41 -07:00
parent 3d5cc0ada4
commit c258fc1590
10 changed files with 376 additions and 190 deletions

View File

@@ -0,0 +1,7 @@
---
"@runfusion/fusion": patch
---
summary: Make Grok CLI chat replies reliable by using the stable headless JSON response.
category: fix
dev: Grok runtime now invokes `grok -p <prompt> --output-format json` and diagnoses empty non-EndTurn results.

View File

@@ -1,10 +1,10 @@
# Grok CLI Contract (FN-7790) # Grok CLI Contract (FN-7790, updated by FN-7796)
Date: 2026-07-10 Date: 2026-07-10
<!-- <!--
FNXC:GrokCli 2026-07-10-11:17: FNXC:GrokCli 2026-07-10-12:58:
FN-7790 supersedes the FN-7722/FN-7724 contract that targeted `superagent-ai/grok-cli`. Operators actually run xAI's official Grok Build TUI (`grok 0.2.93`), so Fusion must invoke `grok -p <prompt> --output-format streaming-json` and parse `thought`/`text`/`end` events; the old `--prompt`/`--format json` plus `step_*` schema is a wrong-product assumption that produced no assistant response. FN-7796 supersedes FN-7790's streaming assumption. Operators run xAI's official Grok Build TUI (`grok 0.2.93`); its `--output-format streaming-json` path intermittently ends `stopReason:"Cancelled"` with zero `text` events, so Fusion's reliable headless prompt path invokes `grok -p <prompt> --output-format json` and parses the single `{text,stopReason,sessionId,requestId,thought}` object. A non-`EndTurn` stop reason with empty text is a concrete diagnostic, never a silent no-message response.
--> -->
## Ground truth ## Ground truth
@@ -21,7 +21,9 @@ External integration evidence:
The previously documented https://github.com/superagent-ai/grok-cli contract is a different product that happens to use the same binary name. Its `grok --prompt <text> --format json` invocation is not accepted by xAI's CLI. The previously documented https://github.com/superagent-ai/grok-cli contract is a different product that happens to use the same binary name. Its `grok --prompt <text> --format json` invocation is not accepted by xAI's CLI.
## Failure that caused FN-7790 ## Failures that shaped the contract
### Wrong-product flags (FN-7790)
The old adapter invocation fails against the real xAI binary: The old adapter invocation fails against the real xAI binary:
@@ -42,28 +44,77 @@ error: unexpected argument '--prompt' found
Usage: grok --prompt-file <PATH> [PROMPT] Usage: grok --prompt-file <PATH> [PROMPT]
``` ```
Because no NDJSON `text` event is produced, Fusion surfaced a blank/no-message assistant response. Because no renderable assistant text is produced, Fusion surfaced a blank/no-message assistant response.
## Confirmed non-interactive invocation ### Streaming JSON cancellation with zero text (FN-7796)
Use xAI Grok Build TUI's single-turn prompt mode with streaming JSON: FN-7790 correctly switched to xAI's real flags and streaming event union, but live triage found `--output-format streaming-json` is intermittently unreliable. The same authenticated `grok 0.2.93` binary sometimes emits only reasoning events, then ends with `stopReason:"Cancelled"` and no `text` event while still exiting 0 with empty stderr.
Live-captured shape:
```jsonl
{"type":"thought","data":"..."}
{"type":"thought","data":"..."}
{"type":"end","stopReason":"Cancelled","sessionId":"...","requestId":"..."}
```
The adapter previously saw parsed events and a successful close, accumulated empty assistant text, set no error, and produced a silent no-message bubble. The reliable replacement is the single-object JSON contract below.
## Confirmed non-interactive invocation used by Fusion
Use xAI Grok Build TUI's single-turn prompt mode with **single-object JSON**:
```bash ```bash
grok -p "<text>" --output-format streaming-json grok -p "<text>" --output-format json
# equivalent long prompt flag: # equivalent long prompt flag:
grok --single "<text>" --output-format streaming-json grok --single "<text>" --output-format json
``` ```
Supported companion flags used by Fusion: Supported companion flags used by Fusion:
- `-p, --single <PROMPT>` — run a single prompt, print the response, and exit. This does not require interactive stdin. - `-p, --single <PROMPT>` — run a single prompt, print the response, and exit. This does not require interactive stdin.
- `--output-format <plain|json|streaming-json>` — streaming adapter uses `streaming-json`. - `--output-format <plain|json|streaming-json>` — Fusion uses `json` for reliable headless prompts.
- `-m, --model <MODEL>` — optional concrete model id. Fusion omits this for the model-less `grok/default` Runtime-mode path. - `-m, --model <MODEL>` — optional concrete model id. Fusion omits this for the model-less `grok/default` Runtime-mode path.
- `--cwd <CWD>` — optional working directory. This replaces the wrong-product `--directory` flag. - `--cwd <CWD>` — optional working directory. This replaces the wrong-product `--directory` flag.
Other observed flags include `--prompt-file <PATH>`, `--prompt-json <JSON>`, `-s/--session-id <UUID>`, `--sandbox <PROFILE>`, `--system-prompt-override <PROMPT>`, and `--max-turns <N>`, but Fusion's adapter does not currently use them. Other observed flags include `--prompt-file <PATH>`, `--prompt-json <JSON>`, `-s/--session-id <UUID>`, `--sandbox <PROFILE>`, `--system-prompt-override <PROMPT>`, and `--max-turns <N>`, but Fusion's adapter does not currently use them.
## Streaming JSON event schema ## Reliable JSON response schema
`--output-format json` emits one final JSON object rather than an NDJSON stream. Observed shape:
```ts
interface GrokJsonResponse {
text?: string;
stopReason?: string;
sessionId?: string;
requestId?: string;
thought?: string;
}
```
Example:
```json
{
"text": "Hello",
"stopReason": "EndTurn",
"sessionId": "019f4d81-8fb1-7f11-98ca-5ae00654b518",
"requestId": "bb7952e2-f1bc-4574-b409-5cc568817fe5",
"thought": "The user wants me to say hello in one word..."
}
```
Mapping in Fusion:
- `thought` → `onThinking(thought)` when non-empty.
- `text` → `onText(text)` and accumulated assistant content when non-empty.
- `sessionId` → `session.sessionId` when present.
- subprocess `close` remains the authoritative promise resolution point because it carries exit status/stderr diagnostics.
Live reliability evidence from FN-7796: `grok -p "say hello in one word" --output-format json` returned real text with `stopReason:"EndTurn"` on 4/4 direct runs, and the built `GrokRuntimeAdapter` carried real text through `onText`/persisted assistant content on 3/3 end-to-end runs against the real binary.
## Streaming JSON event schema (not the primary prompt path)
`--output-format streaming-json` emits one JSON object per line: `--output-format streaming-json` emits one JSON object per line:
@@ -74,13 +125,7 @@ type GrokStreamingJsonEvent =
| { type: "end"; stopReason?: string; sessionId?: string; requestId?: string }; | { type: "end"; stopReason?: string; sessionId?: string; requestId?: string };
``` ```
Mapping in Fusion: Successful captured tail:
- `thought.data` → `onThinking(thought.data)`.
- `text.data` → `onText(text.data)` and accumulated assistant content.
- `end.sessionId` → `session.sessionId` when present. `end` pre-signals terminal output, but subprocess `close` remains the authoritative promise resolution point because it carries exit status/stderr diagnostics.
Real captured tail:
```jsonl ```jsonl
{"type":"thought","data":" one"} {"type":"thought","data":" one"}
@@ -93,25 +138,11 @@ Real captured tail:
{"type":"end","stopReason":"EndTurn","sessionId":"019f4d1e-2582-70e0-a174-c8774782ab01","requestId":"2233f1dc-e9ad-4ae4-8221-caa6afade07f"} {"type":"end","stopReason":"EndTurn","sessionId":"019f4d1e-2582-70e0-a174-c8774782ab01","requestId":"2233f1dc-e9ad-4ae4-8221-caa6afade07f"}
``` ```
A successful run exits 0 with empty stderr. Fusion does not use streaming-json as the primary headless prompt path because it intermittently produces the cancelled/no-text shape documented above. Parser support remains only to keep diagnostics and regression tests concrete if captured streaming output appears in buffered stdout.
## Non-streaming formats ## Other output formats
`--output-format plain` prints renderable response text. `--output-format plain` prints renderable response text, but does not expose `sessionId`, `requestId`, `stopReason`, or `thought`.
`--output-format json` emits one final JSON object rather than an NDJSON stream. Observed shape:
```json
{
"text": "hi",
"stopReason": "EndTurn",
"sessionId": "019f4d18-875b-7662-9bc5-9b71fa0aa6b0",
"requestId": "0e8ef53f-5a5f-4564-a8fd-0200ef96440e",
"thought": "The user wants me to say hi in one word..."
}
```
Fusion uses `streaming-json` for live `onText`/`onThinking` callbacks.
## Model discovery ## Model discovery
@@ -131,7 +162,7 @@ Fusion parses the bullet list conservatively and exposes ids under provider `gro
## Auth and readiness ## Auth and readiness
The CLI owns authentication for CLI-routed execution. Fusion's readiness probe uses `grok --version`; a passing probe proves only that a compatible-looking binary exists, not that the prompt path is authenticated or serviceable. The prompt path is proven by a real `grok -p ... --output-format streaming-json` run. The CLI owns authentication for CLI-routed execution. Fusion's readiness probe uses `grok --version`; a passing probe proves only that a compatible-looking binary exists, not that the prompt path is authenticated or serviceable. The prompt path is proven by a real `grok -p ... --output-format json` run.
Fusion-visible `GROK_API_KEY` remains relevant for the direct xAI OpenAI-compatible endpoint. For CLI-routed sessions, Fusion does not need to see a key as long as the operator-installed CLI is authenticated by its own supported mechanism. Fusion-visible `GROK_API_KEY` remains relevant for the direct xAI OpenAI-compatible endpoint. For CLI-routed sessions, Fusion does not need to see a key as long as the operator-installed CLI is authenticated by its own supported mechanism.
@@ -150,8 +181,9 @@ The adapter preserves the resolve-never-reject runtime contract while surfacing
- spawn failure → `session.state.errorMessage` and diagnostic `onText`. - spawn failure → `session.state.errorMessage` and diagnostic `onText`.
- non-zero subprocess close with no text → stderr/exit diagnostic. - non-zero subprocess close with no text → stderr/exit diagnostic.
- code-0 close with zero parsed NDJSON → wrong-binary/interactive-EOF diagnostic. - code-0 close with no parseable JSON response → wrong-binary/interactive-EOF diagnostic.
- parsed `end` with no accumulated assistant text → legitimate silent response, not a diagnostic. - parseable response with no text and `stopReason !== "EndTurn"` → stop-reason diagnostic, e.g. `Grok CLI ended with stopReason Cancelled and produced no assistant text.`
- parseable `EndTurn` response with no assistant text → legitimate silent response, not a diagnostic.
- text emitted before a noisy/non-zero close → keep the assistant text and avoid replacing it with an error. - text emitted before a noisy/non-zero close → keep the assistant text and avoid replacing it with an error.
This invariant prevents the original blank/no-message symptom while still allowing genuinely empty model turns. This invariant prevents the original blank/no-message symptom while still allowing genuinely empty model turns.

View File

@@ -23,23 +23,23 @@ The previously assumed `superagent-ai/grok-cli` contract is a different product
- **Auth model — the `grok` CLI owns its own authentication; Fusion does not require a Fusion-visible API key to enable/use it (FN-7716).** Fusion additionally probes the `GROK_API_KEY` env var and `~/.grok/user-settings.json` → `{ "apiKey": "..." }` purely as a **non-blocking informational hint** (`apiKeyDetected`); it never gates Enable or the authenticated state. The direct xAI OpenAI-compatible streaming path (base URL `https://api.x.ai/v1`) still uses `$GROK_API_KEY` when present, independent of the CLI provider. - **Auth model — the `grok` CLI owns its own authentication; Fusion does not require a Fusion-visible API key to enable/use it (FN-7716).** Fusion additionally probes the `GROK_API_KEY` env var and `~/.grok/user-settings.json` → `{ "apiKey": "..." }` purely as a **non-blocking informational hint** (`apiKeyDetected`); it never gates Enable or the authenticated state. The direct xAI OpenAI-compatible streaming path (base URL `https://api.x.ai/v1`) still uses `$GROK_API_KEY` when present, independent of the CLI provider.
- Model discovery: `grok models` (plain text). The observed xAI shape is `Default model: <id>`, then `Available models:`, then `* <id> (default)` / `- <id>` bullet rows. - Model discovery: `grok models` (plain text). The observed xAI shape is `Default model: <id>`, then `Available models:`, then `* <id> (default)` / `- <id>` bullet rows.
## CLI streaming execution path (FN-7790) ## CLI headless execution path (FN-7790 / FN-7796)
The plugin's `GrokRuntimeAdapter` streams a real Grok response through xAI's CLI: The plugin's `GrokRuntimeAdapter` returns a real Grok response through xAI's reliable single-object CLI output:
```bash ```bash
grok -p "<text>" --output-format streaming-json grok -p "<text>" --output-format json
# with optional model/cwd: # with optional model/cwd:
grok -p "<text>" --output-format streaming-json -m "grok-4.5" --cwd "/path/to/project" grok -p "<text>" --output-format json -m "grok-4.5" --cwd "/path/to/project"
``` ```
- `-p, --single <PROMPT>` runs a single prompt and exits; it does not require interactive stdin. - `-p, --single <PROMPT>` runs a single prompt and exits; it does not require interactive stdin.
- `--output-format streaming-json` emits NDJSON with event types `thought`, `text`, and `end`. - `--output-format json` emits one object with `{ text, stopReason, sessionId, requestId, thought }`. Fusion buffers stdout until subprocess `close`, then bridges `thought` to `onThinking`, `text` to `onText` and persisted assistant content, and stores `sessionId` when present.
- `thought.data` drives `onThinking`; `text.data` drives `onText` and persisted assistant content; `end.sessionId` is stored when present. The subprocess `close` event remains the authoritative resolution point so stderr/exit diagnostics are preserved. - xAI's `--output-format streaming-json` mode is not used for the primary headless path because live `grok 0.2.93` testing found it can intermittently end `stopReason:"Cancelled"` with zero `text` events. A non-`EndTurn` stop reason with empty text now surfaces a concrete diagnostic instead of a blank assistant response; a parseable `EndTurn` with empty text remains a legitimate silent response.
- A wrong-binary/wrong-flag run that emits no parsed NDJSON surfaces a concrete diagnostic instead of a blank assistant response. A real `end` event with empty accumulated text remains a legitimate silent response. - A wrong-binary/wrong-flag run that emits no parseable JSON surfaces a concrete diagnostic instead of a blank assistant response.
- **Auth implication:** because the `grok` binary resolves its own credentials for this path, a CLI-routed selection needs **no Fusion-visible `GROK_API_KEY`** — unlike the direct xAI OpenAI-compatible streaming path. - **Auth implication:** because the `grok` binary resolves its own credentials for this path, a CLI-routed selection needs **no Fusion-visible `GROK_API_KEY`** — unlike the direct xAI OpenAI-compatible streaming path.
See `docs/grok-cli-contract.md` for the full contract, live captures, and the reason Fusion no longer uses the old `grok --prompt <text> --format json` / `step_*` schema. See `docs/grok-cli-contract.md` for the full contract, live captures, and the reason Fusion no longer uses the old `grok --prompt <text> --format json` / `step_*` schema or the flaky streaming-json prompt path.
## Routing Grok through the CLI runtime (FN-7725 / FN-7753 / FN-7790) ## Routing Grok through the CLI runtime (FN-7725 / FN-7753 / FN-7790)

View File

@@ -42,7 +42,7 @@ describe("spawnGrokStream", () => {
"-p", "-p",
"hello", "hello",
"--output-format", "--output-format",
"streaming-json", "json",
"-m", "-m",
"grok-4.5", "grok-4.5",
"--cwd", "--cwd",
@@ -62,7 +62,7 @@ describe("spawnGrokStream", () => {
"-p", "-p",
"hello", "hello",
"--output-format", "--output-format",
"streaming-json", "json",
"--cwd", "--cwd",
"/tmp/project", "/tmp/project",
], expect.objectContaining({ cwd: "/tmp/project" })); ], expect.objectContaining({ cwd: "/tmp/project" }));

View File

@@ -5,8 +5,8 @@ import type { GrokStreamProcess } from "../cli-stream.js";
import { GrokRuntimeAdapter } from "../runtime-adapter.js"; import { GrokRuntimeAdapter } from "../runtime-adapter.js";
/* /*
FNXC:GrokCli 2026-07-10-11:05: FNXC:GrokCli 2026-07-10-12:54:
FN-7790: adapter tests are pinned to the operator-verified xAI Grok Build TUI stream (`thought`/`text`/`end` with `data`). They intentionally avoid a live binary in CI but exercise the same spawn seam and lifecycle diagnostics that previously hid the wrong `--prompt`/`--format json` contract behind fake superagent-ai fixtures. FN-7796: adapter tests are pinned to the reliable xAI Grok Build TUI headless contract (`--output-format json` single object) and the live-captured flaky `streaming-json` cancellation shape. They intentionally avoid a live binary in CI but exercise the same spawn seam and lifecycle diagnostics that previously hid wrong-contract and cancelled-no-text failures behind fake fixtures.
*/ */
function makeFakeProc(): { proc: GrokStreamProcess; stdout: PassThrough; stderr: PassThrough; kill: ReturnType<typeof vi.fn> } { function makeFakeProc(): { proc: GrokStreamProcess; stdout: PassThrough; stderr: PassThrough; kill: ReturnType<typeof vi.fn> } {
@@ -58,6 +58,65 @@ describe("GrokRuntimeAdapter", () => {
expect(spawn).toHaveBeenCalledWith("grok", "hello grok", expect.objectContaining({ model: undefined })); expect(spawn).toHaveBeenCalledWith("grok", "hello grok", expect.objectContaining({ model: undefined }));
}); });
it("bridges the reliable single-object json response and persists assistant content", async () => {
const { proc, stdout } = makeFakeProc();
const spawn = vi.fn().mockReturnValue(proc);
const adapter = new GrokRuntimeAdapter({ spawn });
const onText = vi.fn();
const onThinking = vi.fn();
const { session } = await adapter.createSession({ onText, onThinking });
const promise = adapter.promptWithFallback(session, "hello grok");
stdout.write(JSON.stringify({ text: "Hello", stopReason: "EndTurn", sessionId: "session-json", requestId: "request-json", thought: "Thinking" }));
stdout.end();
closeProc(proc);
await promise;
expect(onThinking).toHaveBeenCalledWith("Thinking");
expect(onText).toHaveBeenCalledWith("Hello");
expect(session.sessionId).toBe("session-json");
expect(session.state.messages).toContainEqual({ role: "assistant", content: "Hello" });
});
it("surfaces cancelled no-text json object as a diagnostic instead of a silent empty response", async () => {
const { proc, stdout } = makeFakeProc();
const spawn = vi.fn().mockReturnValue(proc);
const adapter = new GrokRuntimeAdapter({ spawn });
const onText = vi.fn();
const { session } = await adapter.createSession({ onText });
const promise = adapter.promptWithFallback(session, "say hello in one word");
stdout.write(JSON.stringify({ text: "", stopReason: "Cancelled", sessionId: "session-cancelled" }));
stdout.end();
closeProc(proc);
await promise;
expect(session.state.errorMessage).toBe("Grok CLI ended with stopReason Cancelled and produced no assistant text.");
expect(onText).toHaveBeenCalledWith(session.state.errorMessage);
expect(session.state.messages).toContainEqual({ role: "assistant", content: session.state.errorMessage });
});
it("surfaces cancelled no-text streaming-json shape as a diagnostic instead of a silent empty response", async () => {
const { proc, stdout } = makeFakeProc();
const spawn = vi.fn().mockReturnValue(proc);
const adapter = new GrokRuntimeAdapter({ spawn });
const onText = vi.fn();
const onThinking = vi.fn();
const { session } = await adapter.createSession({ onText, onThinking });
const promise = adapter.promptWithFallback(session, "say hello in one word");
stdout.write(`${JSON.stringify({ type: "thought", data: "Thinking" })}\n`);
stdout.write(`${JSON.stringify({ type: "end", stopReason: "Cancelled", sessionId: "session-cancelled", requestId: "request-cancelled" })}\n`);
stdout.end();
closeProc(proc);
await promise;
expect(session.state.errorMessage).toBe("Grok CLI ended with stopReason Cancelled and produced no assistant text.");
expect(onText).toHaveBeenCalledWith(session.state.errorMessage);
expect(session.state.messages).toContainEqual({ role: "assistant", content: session.state.errorMessage });
});
it("bridges real xAI thought/text/end events and persists assistant content", async () => { it("bridges real xAI thought/text/end events and persists assistant content", async () => {
const { proc, stdout } = makeFakeProc(); const { proc, stdout } = makeFakeProc();
const spawn = vi.fn().mockReturnValue(proc); const spawn = vi.fn().mockReturnValue(proc);
@@ -75,7 +134,7 @@ describe("GrokRuntimeAdapter", () => {
await promise; await promise;
expect(onThinking.mock.calls.map((c) => c[0])).toEqual(["Thinking"]); expect(onThinking.mock.calls.map((c) => c[0])).toEqual(["Thinking"]);
expect(onText.mock.calls.map((c) => c[0])).toEqual(["Hel", "lo"]); expect(onText.mock.calls.map((c) => c[0])).toEqual(["Hello"]);
expect(session.sessionId).toBe("session-1"); expect(session.sessionId).toBe("session-1");
expect(session.state.messages).toContainEqual({ role: "assistant", content: "Hello" }); expect(session.state.messages).toContainEqual({ role: "assistant", content: "Hello" });
}); });
@@ -171,7 +230,7 @@ describe("GrokRuntimeAdapter", () => {
expect(session.state.errorMessage).toBe("Grok CLI failed with code 2 and no stderr output."); expect(session.state.errorMessage).toBe("Grok CLI failed with code 2 and no stderr output.");
}); });
it("records a concrete diagnostic for code-0 exits with zero NDJSON output", async () => { it("records a concrete diagnostic for code-0 exits with zero JSON output", async () => {
const { proc, stdout } = makeFakeProc(); const { proc, stdout } = makeFakeProc();
const spawn = vi.fn().mockReturnValue(proc); const spawn = vi.fn().mockReturnValue(proc);
const adapter = new GrokRuntimeAdapter({ spawn }); const adapter = new GrokRuntimeAdapter({ spawn });
@@ -184,13 +243,13 @@ describe("GrokRuntimeAdapter", () => {
await promise; await promise;
expect(session.state.errorMessage).toBe( expect(session.state.errorMessage).toBe(
"Grok CLI produced no NDJSON output for a headless prompt; this usually means the binary on PATH is not xAI's supported Grok Build TUI headless implementation, did not recognize -p/--output-format streaming-json, or exited interactive mode immediately after stdin EOF.", "Grok CLI produced no JSON output for a headless prompt; this usually means the binary on PATH is not xAI's supported Grok Build TUI headless implementation, did not recognize -p/--output-format json, or exited interactive mode immediately after stdin EOF.",
); );
expect(onText).toHaveBeenCalledWith(session.state.errorMessage); expect(onText).toHaveBeenCalledWith(session.state.errorMessage);
expect(session.state.messages).toContainEqual({ role: "assistant", content: session.state.errorMessage }); expect(session.state.messages).toContainEqual({ role: "assistant", content: session.state.errorMessage });
}); });
it("records a concrete diagnostic for code-0 exits with non-NDJSON stdout only", async () => { it("records a concrete diagnostic for code-0 exits with non-JSON stdout only", async () => {
const { proc, stdout } = makeFakeProc(); const { proc, stdout } = makeFakeProc();
const spawn = vi.fn().mockReturnValue(proc); const spawn = vi.fn().mockReturnValue(proc);
const adapter = new GrokRuntimeAdapter({ spawn }); const adapter = new GrokRuntimeAdapter({ spawn });
@@ -204,7 +263,7 @@ describe("GrokRuntimeAdapter", () => {
await promise; await promise;
expect(session.state.errorMessage).toBe( expect(session.state.errorMessage).toBe(
"Grok CLI produced stdout but no NDJSON events for a headless prompt; first line: Welcome to grok interactive mode", "Grok CLI produced stdout but no parseable JSON response for a headless prompt; first output: Welcome to grok interactive mode",
); );
expect(onText).toHaveBeenCalledWith(session.state.errorMessage); expect(onText).toHaveBeenCalledWith(session.state.errorMessage);
}); });

View File

@@ -1,56 +1,80 @@
import { describe, expect, it } from "vitest"; import { describe, expect, it } from "vitest";
import { parseLine } from "../stream-parser.js"; import { parseJsonOutput, parseLine } from "../stream-parser.js";
/* /*
FNXC:GrokCli 2026-07-10-11:02: FNXC:GrokCli 2026-07-10-12:53:
FN-7790: fixtures use xAI Grok Build TUI's real `--output-format streaming-json` schema captured from the operator binary. Keep tests on `thought`/`text`/`end` so a future wrong-product `step_*`/`tool_use` assumption fails deterministically before production returns no messages again. FN-7796: fixtures pin the reliable xAI Grok Build TUI headless contract, `--output-format json`, because live `streaming-json` intermittently ended `stopReason:"Cancelled"` without text. Keep one streaming parser regression for the captured cancelled shape so diagnostics stay concrete if the flaky shape appears in buffered output.
*/ */
describe("parseLine (xAI Grok CLI streaming-json)", () => { describe("parseJsonOutput (xAI Grok CLI json)", () => {
it("parses a thought event", () => { it("parses the reliable single-object response", () => {
const line = JSON.stringify({ type: "thought", data: "Thinking" }); const output = JSON.stringify({
expect(parseLine(line)).toEqual({ type: "thought", data: "Thinking" }); text: "Hello",
stopReason: "EndTurn",
sessionId: "session-1",
requestId: "request-1",
thought: "Thinking",
});
expect(parseJsonOutput(output)).toEqual({
text: "Hello",
stopReason: "EndTurn",
sessionId: "session-1",
requestId: "request-1",
thought: "Thinking",
});
}); });
it("parses a text event", () => { it("tolerates pretty-printed json from the real CLI", () => {
const line = JSON.stringify({ type: "text", data: "Hello" }); const output = `\n{\n "text": "Hello",\n "stopReason": "EndTurn",\n "sessionId": "session-1",\n "requestId": "request-1",\n "thought": "Thinking"\n}\n`;
expect(parseLine(line)).toEqual({ type: "text", data: "Hello" }); expect(parseJsonOutput(output)?.text).toBe("Hello");
}); });
it("parses an end event", () => { it("preserves a terminal empty EndTurn object", () => {
expect(parseJsonOutput(JSON.stringify({ text: "", stopReason: "EndTurn" }))).toEqual({
text: "",
stopReason: "EndTurn",
sessionId: undefined,
requestId: undefined,
thought: undefined,
});
});
it("skips empty, non-JSON, malformed JSON, arrays, and unrelated objects without throwing", () => {
expect(parseJsonOutput("")).toBeNull();
expect(parseJsonOutput("Welcome to grok interactive mode")).toBeNull();
expect(() => parseJsonOutput("{not valid json")).not.toThrow();
expect(parseJsonOutput("{not valid json")).toBeNull();
expect(parseJsonOutput(JSON.stringify([{ text: "hi" }]))).toBeNull();
expect(parseJsonOutput(JSON.stringify({ type: "step_start" }))).toBeNull();
});
});
describe("parseLine (captured flaky streaming-json diagnostics)", () => {
it("parses the cancelled no-text terminal shape", () => {
const line = JSON.stringify({ const line = JSON.stringify({
type: "end", type: "end",
stopReason: "EndTurn", stopReason: "Cancelled",
sessionId: "session-1", sessionId: "session-1",
requestId: "request-1", requestId: "request-1",
}); });
expect(parseLine(line)).toEqual({ expect(parseLine(line)).toEqual({
type: "end", type: "end",
stopReason: "EndTurn", stopReason: "Cancelled",
sessionId: "session-1", sessionId: "session-1",
requestId: "request-1", requestId: "request-1",
}); });
}); });
it("skips empty and non-JSON lines", () => { it("parses thought/text events for buffered streaming regressions", () => {
expect(parseLine(JSON.stringify({ type: "thought", data: "Thinking" }))).toEqual({ type: "thought", data: "Thinking" });
expect(parseLine(JSON.stringify({ type: "text", data: "Hello" }))).toEqual({ type: "text", data: "Hello" });
});
it("skips malformed, unknown, and legacy wrong-product lines", () => {
expect(parseLine("")).toBeNull(); expect(parseLine("")).toBeNull();
expect(parseLine(" ")).toBeNull();
expect(parseLine("[SandboxDebug] booting")).toBeNull(); expect(parseLine("[SandboxDebug] booting")).toBeNull();
});
it("skips malformed JSON without throwing", () => {
expect(() => parseLine("{not valid json")).not.toThrow();
expect(parseLine("{not valid json")).toBeNull(); expect(parseLine("{not valid json")).toBeNull();
});
it("skips missing, unknown, and legacy wrong-product event types", () => {
expect(parseLine(JSON.stringify({ foo: "bar" }))).toBeNull();
expect(parseLine(JSON.stringify({ type: "some_future_event", data: 1 }))).toBeNull();
expect(parseLine(JSON.stringify({ type: "step_start", stepNumber: 1 }))).toBeNull(); expect(parseLine(JSON.stringify({ type: "step_start", stepNumber: 1 }))).toBeNull();
expect(parseLine(JSON.stringify({ type: "tool_use", toolCall: {}, toolResult: {} }))).toBeNull();
});
it("skips a JSON array", () => {
expect(parseLine(JSON.stringify([{ type: "text", data: "hi" }]))).toBeNull(); expect(parseLine(JSON.stringify([{ type: "text", data: "hi" }]))).toBeNull();
}); });
}); });

View File

@@ -2,8 +2,8 @@ import { spawn, type ChildProcessByStdio } from "node:child_process";
import type { Readable } from "node:stream"; import type { Readable } from "node:stream";
/* /*
FNXC:GrokCli 2026-07-10-10:49: FNXC:GrokCli 2026-07-10-12:50:
FN-7790: the operator-installed binary is xAI's Grok Build TUI (`grok 0.2.93`), not the previously assumed `superagent-ai/grok-cli`. Invoke the real headless contract, `grok -p <prompt> --output-format streaming-json [-m <model>] [--cwd <dir>]`; the old `--prompt`/`--format json`/`--directory` flags are rejected by the real binary and produce zero assistant text. Keep the existing foreground pipe and Windows shell handling so the adapter can stream line-by-line NDJSON without raw detached processes. FN-7796: the operator-installed binary is xAI's Grok Build TUI (`grok 0.2.93`). Its `--output-format streaming-json` mode intermittently ends `stopReason:"Cancelled"` with zero `text` events, so Fusion's headless prompt path uses the reliable single-object contract, `grok -p <prompt> --output-format json [-m <model>] [--cwd <dir>]`. Keep foreground piped stdio and Windows shell handling so the adapter can buffer stdout, parse the object on close, and surface close/stderr diagnostics without raw detached processes.
*/ */
export type GrokStreamProcess = ChildProcessByStdio<null, Readable, Readable>; export type GrokStreamProcess = ChildProcessByStdio<null, Readable, Readable>;
@@ -15,14 +15,12 @@ export interface SpawnGrokStreamOptions {
} }
/** /**
* Spawn `grok -p <prompt> --output-format streaming-json [-m <model>] [--cwd <cwd>]` * Spawn `grok -p <prompt> --output-format json [-m <model>] [--cwd <cwd>]`
* with piped stdio for line-by-line NDJSON consumption via readline. * with piped stdio. The adapter buffers stdout and parses the complete
* * single-object response on subprocess close.
* Does not read/buffer output itself — callers attach a `readline` interface
* to `proc.stdout` (see `runtime-adapter.ts`).
*/ */
export function spawnGrokStream(binary: string, prompt: string, options?: SpawnGrokStreamOptions): GrokStreamProcess { export function spawnGrokStream(binary: string, prompt: string, options?: SpawnGrokStreamOptions): GrokStreamProcess {
const args: string[] = ["-p", prompt, "--output-format", "streaming-json"]; const args: string[] = ["-p", prompt, "--output-format", "json"];
const model = options?.model?.trim(); const model = options?.model?.trim();
if (model) { if (model) {
// FNXC:GrokCliRouting 2026-07-10-10:49: FN-7790 keeps FN-7753's concrete `grok-cli/*` model preservation but uses xAI Grok Build TUI's accepted short flag, `-m <model>`, with the provider prefix stripped by runtime-adapter.ts. // FNXC:GrokCliRouting 2026-07-10-10:49: FN-7790 keeps FN-7753's concrete `grok-cli/*` model preservation but uses xAI Grok Build TUI's accepted short flag, `-m <model>`, with the provider prefix stripped by runtime-adapter.ts.

View File

@@ -1,27 +1,26 @@
import { createInterface } from "node:readline";
import { forceKillGrokStream, spawnGrokStream, type GrokStreamProcess, type SpawnGrokStreamOptions } from "./cli-stream.js"; import { forceKillGrokStream, spawnGrokStream, type GrokStreamProcess, type SpawnGrokStreamOptions } from "./cli-stream.js";
import { parseLine } from "./stream-parser.js"; import { parseJsonOutput, parseLine } from "./stream-parser.js";
import type { AgentRuntime, AgentRuntimeOptions, AgentSession, AgentSessionResult, GrokSession } from "./types.js"; import type { AgentRuntime, AgentRuntimeOptions, AgentSession, AgentSessionResult, GrokSession } from "./types.js";
/* /*
FNXC:GrokCli 2026-07-10-10:54: FNXC:GrokCli 2026-07-10-12:52:
FN-7790: the production binary is xAI's Grok Build TUI, whose non-interactive prompt path is `grok -p <text> --output-format streaming-json` and whose NDJSON union is `thought`/`text`/`end` with payloads in `data`. Bridge `text.data` to `onText`, `thought.data` to `onThinking`, and record `end.sessionId` without resolving before subprocess close, because close still carries stderr/exit diagnostics. The obsolete `step_*`/`tool_use`/`error` handling targeted a different `grok` product and is intentionally removed so tests cannot pass on the wrong schema again. FN-7796: the production binary is xAI's Grok Build TUI. Its `--output-format streaming-json` path intermittently emits only `thought` events and then `stopReason:"Cancelled"` with no `text`, so the adapter now consumes the reliable `--output-format json` single object on subprocess close. Bridge object `text` to `onText`, object `thought` to `onThinking`, record `sessionId`, and make non-`EndTurn` empty-text terminals diagnosable instead of silent.
FNXC:GrokCliRouting 2026-07-10-10:54: FNXC:GrokCliRouting 2026-07-10-10:54:
FN-7753's auto-derived `grok` runtime routing from a `grok-cli/*` model selection still preserves the concrete model. Normalize provider-qualified ids (`grok-cli/<id>` or `grok/<id>`) at session creation/prompt time and pass only the concrete id to `grok -m`; the no-model Runtime-mode path keeps the historical `grok/default` session fallback and omits `-m`. FN-7753's auto-derived `grok` runtime routing from a `grok-cli/*` model selection still preserves the concrete model. Normalize provider-qualified ids (`grok-cli/<id>` or `grok/<id>`) at session creation/prompt time and pass only the concrete id to `grok -m`; the no-model Runtime-mode path keeps the historical `grok/default` session fallback and omits `-m`.
*/ */
/** /**
* Cold-start ceiling: if `grok -p --output-format streaming-json` produces no * Cold-start ceiling: if `grok -p --output-format json` produces no stdout
* stdout line within this window, treat it as a hung/failed subprocess and * bytes within this window, treat it as a hung/failed subprocess and resolve
* resolve (never reject — mirrors the Droid adapter's resolve-on-error lifecycle * (never reject — mirrors the Droid adapter's resolve-on-error lifecycle so pi
* so pi always gets a well-formed, if empty, result instead of an unhandled rejection). * always gets a well-formed, if diagnostic, result instead of an unhandled rejection).
*/ */
const FIRST_LINE_TIMEOUT_MS = 60_000; const FIRST_OUTPUT_TIMEOUT_MS = 60_000;
/** /**
* Inactivity safety net: kill the subprocess if no stdout line arrives for * Inactivity safety net: kill the subprocess if no stdout bytes arrive for
* this long after the first line. Generous ceiling mirroring the Droid * this long after the first chunk. Generous ceiling mirroring the Droid
* adapter's rationale — the caller (Fusion's stuck-task detection / abort * adapter's rationale — the caller (Fusion's stuck-task detection / abort
* signal) is the authoritative "this session is stuck" source; this is a * signal) is the authoritative "this session is stuck" source; this is a
* last-resort guard for a catastrophically hung `grok` process. * last-resort guard for a catastrophically hung `grok` process.
@@ -55,18 +54,64 @@ function formatCloseDiagnostic(code: number | null, signal: NodeJS.Signals | nul
return detail ? `Grok CLI failed (${exitDetail}): ${detail}` : `Grok CLI failed with ${exitDetail} and no stderr output.`; return detail ? `Grok CLI failed (${exitDetail}): ${detail}` : `Grok CLI failed with ${exitDetail} and no stderr output.`;
} }
function formatNoNdjsonDiagnostic(firstStdoutLine: string | undefined): string { function formatNoJsonDiagnostic(firstStdoutChunk: string | undefined): string {
const firstLine = firstStdoutLine ? compactDiagnostic(firstStdoutLine) : ""; const firstChunk = firstStdoutChunk ? compactDiagnostic(firstStdoutChunk) : "";
if (firstLine) { if (firstChunk) {
return `Grok CLI produced stdout but no NDJSON events for a headless prompt; first line: ${firstLine}`; return `Grok CLI produced stdout but no parseable JSON response for a headless prompt; first output: ${firstChunk}`;
} }
return "Grok CLI produced no NDJSON output for a headless prompt; this usually means the binary on PATH is not xAI's supported Grok Build TUI headless implementation, did not recognize -p/--output-format streaming-json, or exited interactive mode immediately after stdin EOF."; return "Grok CLI produced no JSON output for a headless prompt; this usually means the binary on PATH is not xAI's supported Grok Build TUI headless implementation, did not recognize -p/--output-format json, or exited interactive mode immediately after stdin EOF.";
}
function formatTerminalNoTextDiagnostic(stopReason: string): string {
return `Grok CLI ended with stopReason ${stopReason} and produced no assistant text.`;
} }
function appendMessage(session: GrokSession, role: "user" | "assistant", content: string): void { function appendMessage(session: GrokSession, role: "user" | "assistant", content: string): void {
session.state.messages.push({ role, content }); session.state.messages.push({ role, content });
} }
interface ParsedPromptOutput {
text: string;
thought?: string;
stopReason?: string;
sessionId?: string;
parsed: boolean;
}
function parsePromptOutput(stdout: string): ParsedPromptOutput {
const json = parseJsonOutput(stdout);
if (json) {
return {
text: json.text ?? "",
thought: json.thought,
stopReason: json.stopReason,
sessionId: json.sessionId,
parsed: true,
};
}
let text = "";
let thought = "";
let stopReason: string | undefined;
let sessionId: string | undefined;
let parsed = false;
for (const line of stdout.split(/\r?\n/)) {
const event = parseLine(line);
if (!event) continue;
parsed = true;
if (event.type === "text") {
text += event.data;
} else if (event.type === "thought") {
thought += event.data;
} else {
stopReason = event.stopReason;
sessionId = event.sessionId;
}
}
return { text, thought: thought || undefined, stopReason, sessionId, parsed };
}
export interface GrokRuntimeAdapterOptions { export interface GrokRuntimeAdapterOptions {
/** Binary name/path to invoke. Defaults to "grok" (PATH resolution). */ /** Binary name/path to invoke. Defaults to "grok" (PATH resolution). */
binary?: string; binary?: string;
@@ -125,10 +170,6 @@ export class GrokRuntimeAdapter implements AgentRuntime {
try { try {
proc = this.spawnFn(this.binary, prompt, { cwd, model: modelForCli(grokSession.model), signal }); proc = this.spawnFn(this.binary, prompt, { cwd, model: modelForCli(grokSession.model), signal });
} catch (err) { } catch (err) {
// Spawn threw synchronously (e.g. binary not found without shell
// resolution) — resolve, never reject, matching the CLI-adapter
// contract of always producing a well-formed result while retaining
// the concrete diagnostic for callers that surface session.state.
const message = err instanceof Error ? err.message : String(err); const message = err instanceof Error ? err.message : String(err);
const diagnostic = compactDiagnostic(`Grok CLI spawn failed: ${message}`); const diagnostic = compactDiagnostic(`Grok CLI spawn failed: ${message}`);
grokSession.state.errorMessage = diagnostic; grokSession.state.errorMessage = diagnostic;
@@ -139,14 +180,13 @@ export class GrokRuntimeAdapter implements AgentRuntime {
} }
let settled = false; let settled = false;
let firstLineReceived = false; let firstOutputReceived = false;
let receivedNdjsonEvent = false; let firstStdoutChunk: string | undefined;
let firstStdoutLine: string | undefined;
let receivedText = false;
let assistantText = ""; let assistantText = "";
let diagnosticEmitted = false; let diagnosticEmitted = false;
let stderr = ""; let stderr = "";
let firstLineTimer: NodeJS.Timeout | undefined; let stdout = "";
let firstOutputTimer: NodeJS.Timeout | undefined;
let inactivityTimer: NodeJS.Timeout | undefined; let inactivityTimer: NodeJS.Timeout | undefined;
const setErrorMessage = (message: string) => { const setErrorMessage = (message: string) => {
@@ -156,23 +196,33 @@ export class GrokRuntimeAdapter implements AgentRuntime {
const emitDiagnosticText = (message: string | undefined) => { const emitDiagnosticText = (message: string | undefined) => {
const diagnostic = message?.trim(); const diagnostic = message?.trim();
if (!diagnostic || receivedText || diagnosticEmitted) return; if (!diagnostic || assistantText || diagnosticEmitted) return;
diagnosticEmitted = true; diagnosticEmitted = true;
grokSession.callbacks.onText?.(diagnostic); grokSession.callbacks.onText?.(diagnostic);
appendMessage(grokSession, "assistant", diagnostic); appendMessage(grokSession, "assistant", diagnostic);
}; };
/* const emitParsedOutput = (parsed: ParsedPromptOutput) => {
FNXC:GrokCli 2026-07-10-00:00: if (parsed.thought) {
A failing headless `grok` run can close stdout before the child `close` event reports its non-zero exit and stderr. Resolving on readline close made dashboard Chat persist an empty assistant message before the diagnostic existed. Finalize only from subprocess close/error or lifecycle timeouts, and store concrete stderr/NDJSON error details on session.state.errorMessage so shared chat/executor seams can surface the reason without breaking the resolve-never-reject runtime contract. grokSession.callbacks.onThinking?.(parsed.thought);
}
if (parsed.sessionId) {
grokSession.sessionId = parsed.sessionId;
}
if (parsed.text.length > 0) {
assistantText += parsed.text;
grokSession.callbacks.onText?.(parsed.text);
return;
}
if (parsed.stopReason && parsed.stopReason !== "EndTurn") {
setErrorMessage(formatTerminalNoTextDiagnostic(parsed.stopReason));
}
};
FNXC:GrokCli 2026-07-10-10:56:
FN-7790 keeps FN-7788's zero-output diagnostic but updates the invariant for xAI Grok Build TUI: a valid `grok -p <text> --output-format streaming-json` run emits at least an `end` event, with optional `thought`/`text` events. A code-0 close with zero parsed NDJSON is a wrong-binary/interactive-EOF failure surfaced through both `onText` and `session.state.errorMessage`; a real `end` event with empty assistant text remains a legitimate silent response.
*/
const finish = () => { const finish = () => {
if (settled) return; if (settled) return;
settled = true; settled = true;
if (firstLineTimer) clearTimeout(firstLineTimer); if (firstOutputTimer) clearTimeout(firstOutputTimer);
if (inactivityTimer) clearTimeout(inactivityTimer); if (inactivityTimer) clearTimeout(inactivityTimer);
if (assistantText) { if (assistantText) {
appendMessage(grokSession, "assistant", assistantText); appendMessage(grokSession, "assistant", assistantText);
@@ -193,43 +243,24 @@ export class GrokRuntimeAdapter implements AgentRuntime {
}, INACTIVITY_TIMEOUT_MS); }, INACTIVITY_TIMEOUT_MS);
}; };
firstLineTimer = setTimeout(() => { firstOutputTimer = setTimeout(() => {
if (firstLineReceived) return; if (firstOutputReceived) return;
setErrorMessage( setErrorMessage(
`Grok CLI produced no stdout within ${FIRST_LINE_TIMEOUT_MS}ms for a headless prompt; the process was killed.`, `Grok CLI produced no stdout within ${FIRST_OUTPUT_TIMEOUT_MS}ms for a headless prompt; the process was killed.`,
); );
forceKillGrokStream(proc); forceKillGrokStream(proc);
finish(); finish();
}, FIRST_LINE_TIMEOUT_MS); }, FIRST_OUTPUT_TIMEOUT_MS);
const rl = createInterface({ input: proc.stdout, crlfDelay: Infinity, terminal: false }); proc.stdout?.on("data", (chunk: Buffer | string) => {
const text = chunk.toString();
rl.on("line", (line: string) => { if (!firstOutputReceived) {
if (!firstLineReceived) { firstOutputReceived = true;
firstLineReceived = true; firstStdoutChunk = text;
firstStdoutLine = line; if (firstOutputTimer) clearTimeout(firstOutputTimer);
if (firstLineTimer) clearTimeout(firstLineTimer);
} }
stdout += text;
resetInactivityTimer(); resetInactivityTimer();
const event = parseLine(line);
if (!event) return;
receivedNdjsonEvent = true;
if (event.type === "text") {
if (event.data.length > 0) {
receivedText = true;
assistantText += event.data;
}
grokSession.callbacks.onText?.(event.data);
} else if (event.type === "thought") {
grokSession.callbacks.onThinking?.(event.data);
} else if (event.type === "end") {
grokSession.sessionId = event.sessionId ?? grokSession.sessionId;
}
// `end` is the real xAI stream's terminal marker, but subprocess close
// remains authoritative for resolving because close carries non-zero
// exit/stderr diagnostics for failed runs.
}); });
proc.stderr?.on("data", (chunk: Buffer | string) => { proc.stderr?.on("data", (chunk: Buffer | string) => {
@@ -238,31 +269,26 @@ export class GrokRuntimeAdapter implements AgentRuntime {
proc.on("error", (err) => { proc.on("error", (err) => {
const message = err instanceof Error ? err.message : String(err); const message = err instanceof Error ? err.message : String(err);
if (!receivedText) { if (!assistantText) {
setErrorMessage(compactDiagnostic(`Grok CLI process error: ${message}`)); setErrorMessage(compactDiagnostic(`Grok CLI process error: ${message}`));
} }
finish(); finish();
}); });
proc.on("close", (code: number | null, signal: NodeJS.Signals | null) => { proc.on("close", (code: number | null, signal: NodeJS.Signals | null) => {
try { const parsed = parsePromptOutput(stdout);
rl.close(); if (parsed.parsed) {
} catch { emitParsedOutput(parsed);
// already closed
} }
const failed = typeof code === "number" ? code !== 0 : Boolean(signal); const failed = typeof code === "number" ? code !== 0 : Boolean(signal);
if (!receivedText && failed) { if (!assistantText && failed) {
setErrorMessage(formatCloseDiagnostic(typeof code === "number" ? code : null, signal, stderr)); setErrorMessage(formatCloseDiagnostic(typeof code === "number" ? code : null, signal, stderr));
} else if (!receivedText && !receivedNdjsonEvent && typeof code === "number" && code === 0) { } else if (!assistantText && !parsed.parsed && typeof code === "number" && code === 0) {
setErrorMessage(formatNoNdjsonDiagnostic(firstStdoutLine)); setErrorMessage(formatNoJsonDiagnostic(firstStdoutChunk));
} }
finish(); finish();
}); });
rl.on("close", () => {
// Wait for the child `close` event so non-zero exits can attach stderr
// diagnostics before callers inspect the session.
});
}); });
} }

View File

@@ -1,27 +1,60 @@
import type { GrokNdjsonEvent } from "./types.js"; import type { GrokCliJsonResponse, GrokNdjsonEvent } from "./types.js";
/* /*
FNXC:GrokCli 2026-07-10-10:50: FNXC:GrokCli 2026-07-10-12:50:
FN-7790: xAI's official Grok Build TUI streams newline-delimited `thought`/`text`/`end` JSON from `grok -p <prompt> --output-format streaming-json`. The previously accepted `step_*`/`tool_use`/`error` events described a different `grok` binary and masked production no-message failures, so unknown legacy lines now fall through as unrecognized while the parser keeps its never-throw resilience. FN-7796: xAI Grok Build TUI's `--output-format streaming-json` intermittently ends with `stopReason:"Cancelled"` and zero `text` events. The headless path now uses the reliable single-object `--output-format json` response, so parser callers should parse the complete stdout buffer into `{text,stopReason,sessionId,requestId,thought}` and treat invalid/partial buffers as absent output rather than throwing.
*/ */
const KNOWN_EVENT_TYPES = new Set(["thought", "text", "end"]);
/** /**
* Parse a single NDJSON line from `grok -p --output-format streaming-json` stdout into a * Parse the complete stdout buffer from
* typed event, or null when the line should be skipped (empty, non-JSON * `grok -p <prompt> --output-format json` into the real xAI Grok Build TUI
* debug noise, malformed JSON, or a JSON object whose `type` isn't one of * response object, or null when the output is empty, non-JSON, a JSON array,
* the real xAI streaming event types). * or an unrelated object with none of the expected response fields.
*/ */
export function parseLine(line: string): GrokNdjsonEvent | null { export function parseJsonOutput(output: string): GrokCliJsonResponse | null {
const trimmed = line.trim(); const trimmed = output.trim();
if (!trimmed || !trimmed.startsWith("{")) {
// Skip empty lines return null;
if (!trimmed) { }
return null;
} let parsed: unknown;
try {
// Skip non-JSON lines (e.g. any stray debug/log output not part of the JSONL stream) parsed = JSON.parse(trimmed);
if (!trimmed.startsWith("{")) { } catch {
return null;
}
if (parsed === null || typeof parsed !== "object" || Array.isArray(parsed)) {
return null;
}
const candidate = parsed as Record<string, unknown>;
const hasKnownField = ["text", "stopReason", "sessionId", "requestId", "thought"].some((key) => key in candidate);
if (!hasKnownField) {
return null;
}
return {
text: typeof candidate.text === "string" ? candidate.text : undefined,
stopReason: typeof candidate.stopReason === "string" ? candidate.stopReason : undefined,
sessionId: typeof candidate.sessionId === "string" ? candidate.sessionId : undefined,
requestId: typeof candidate.requestId === "string" ? candidate.requestId : undefined,
thought: typeof candidate.thought === "string" ? candidate.thought : undefined,
};
}
const STREAMING_EVENT_TYPES = new Set(["thought", "text", "end"]);
/**
* Parse a single NDJSON line from the legacy/flaky
* `--output-format streaming-json` contract. The runtime no longer relies on
* this as its primary path, but retaining this parser lets deterministic
* regressions model the live-captured cancelled-no-text stream shape and
* produce a concrete diagnostic instead of treating it as arbitrary garbage.
*/
export function parseLine(line: string): GrokNdjsonEvent | null {
const trimmed = line.trim();
if (!trimmed || !trimmed.startsWith("{")) {
return null; return null;
} }
@@ -29,17 +62,15 @@ export function parseLine(line: string): GrokNdjsonEvent | null {
try { try {
parsed = JSON.parse(trimmed); parsed = JSON.parse(trimmed);
} catch { } catch {
console.error("Failed to parse Grok CLI NDJSON line:", trimmed);
return null; return null;
} }
// Validate that the parsed result is a non-null object (not array, not primitive)
if (parsed === null || typeof parsed !== "object" || Array.isArray(parsed)) { if (parsed === null || typeof parsed !== "object" || Array.isArray(parsed)) {
return null; return null;
} }
const candidate = parsed as { type?: unknown }; const candidate = parsed as { type?: unknown };
if (typeof candidate.type !== "string" || !KNOWN_EVENT_TYPES.has(candidate.type)) { if (typeof candidate.type !== "string" || !STREAMING_EVENT_TYPES.has(candidate.type)) {
return null; return null;
} }

View File

@@ -1,8 +1,17 @@
/* /*
FNXC:GrokCli 2026-07-10-10:48: FNXC:GrokCli 2026-07-10-12:50:
FN-7790: operators run xAI's official Grok Build TUI (`grok 0.2.93`), not the previously assumed `superagent-ai/grok-cli` product. The real headless stream is `grok -p <prompt> --output-format streaming-json` and emits `thought`/`text`/`end` objects with `data`, so these types intentionally retire the old `step_*`/`tool_use`/`error` union that made fake tests pass while the real binary returned no assistant text. FN-7796: xAI Grok Build TUI's `--output-format streaming-json` can emit reasoning-only events and then `stopReason:"Cancelled"` with zero assistant text. The primary headless contract is therefore the reliable single `--output-format json` object `{text,stopReason,sessionId,requestId,thought}`; streaming event types remain only for diagnostics/regressions that model the captured flaky shape.
*/ */
export interface GrokCliJsonResponse {
text?: string;
stopReason?: string;
sessionId?: string;
requestId?: string;
thought?: string;
}
export interface GrokThoughtEvent { export interface GrokThoughtEvent {
type: "thought"; type: "thought";
data: string; data: string;