# Reference

> Every command and flag, file format, and output field.

Everything Spoiler reads and writes, in one place. [Introduction](/docs) explains what each part
is for; [Set up](/docs/set-up) puts them together. `spoiler <command> --help` always lists the
flags of the version you have installed.

## Installing

```sh
curl -fsSL https://spoiler.sh/install | sh
```

The script picks the build for your OS and CPU, checks its SHA-256, and installs it to
`~/.local/bin`. Linux gets a static musl build that runs on any distribution. Pin a release with
`curl -fsSL https://spoiler.sh/install | SPOILER_VERSION=v0.1.0 sh`, and choose the directory
with `SPOILER_INSTALL_DIR`.

| Installer | Command |
| --- | --- |
| uv | `uv tool install spoiler` |
| pip | `pip install spoiler` |
| Cargo (Rust 1.88+) | `cargo install spoiler --locked` |

The Python packages only put the `spoiler` binary on `PATH`; there's no Python API.
[GitHub Releases](https://github.com/sahil-shubham/spoiler/releases) has an archive for each
target, each with a `.sha256` file:

- `aarch64-apple-darwin`, `x86_64-apple-darwin`
- `x86_64-unknown-linux-gnu`, `aarch64-unknown-linux-gnu` (glibc 2.28 or later)
- `x86_64-unknown-linux-musl`, `aarch64-unknown-linux-musl` (static; any Linux)

From a checkout of the repository: `cargo install --path crates/cli --locked`.

## Commands

| Command | Reads | Writes | Network |
| --- | --- | --- | --- |
| `vocab build` | `product.json` and source files | `vocabulary_snapshot` | OpenRouter, unless `--candidate` |
| `vocab check` | A feature map | `vocabulary_check` | None |
| `recordings list` | A PostHog project | `recording_page` | PostHog |
| `recordings fetch` | A PostHog session | `recording` | PostHog |
| `decode` | A recording file | `recording` | None |
| `compile` | A recording and a feature map | `trace` | None |
| `analyze` | A trace, or a recording, and a feature map | `analysis_request` or `analysis` | OpenRouter with `--model` |
| `run` | A recording or PostHog session, and a feature map | `session` | PostHog with `--session`, OpenRouter with `--model` |

Every command:

- writes one JSON artifact, to `--out FILE` or to stdout;
- reads `-` as standard input for any input path, so commands pipe into each other;
- takes all configuration as flags, and credentials only from the environment;
- accepts `--max-input-mib` (default 512), the most a recording may decode to.

```sh
spoiler recordings fetch --project 12345 --session "$SESSION" \
  | spoiler compile --recording - --vocab spoiler/vocab.json --app web \
  | spoiler analyze --trace - --vocab spoiler/vocab.json --prepare-only
```

### vocab build

Builds a feature map from `product.json` and the source files you name.

| Flag | Meaning |
| --- | --- |
| `--config FILE` | The product config. Required. |
| `--source FILE` | A source file for the model to read. Required; repeat for each file. |
| `--model ID` | OpenRouter model to draft the map. Needs `OPENROUTER_API_KEY`. |
| `--candidate FILE` | Package this feature map instead of drafting one; no model call. Takes precedence over `--model`. |
| `--source-revision REV` | The revision the sources were read at, such as a commit, recorded with the map. |
| `--openrouter-url URL` | An OpenRouter-compatible endpoint. Default `https://openrouter.ai/api/v1`. |
| `--timeout SECONDS` | Per request. Default 120. |

### vocab check

`--vocab FILE` validates a feature map, in YAML, JSON, or a built snapshot, and reports its
digest, apps, counts, and `warnings` for entries that can never match. No network.

### recordings list

One page of recordings that started in a time window, oldest first.

| Flag | Meaning |
| --- | --- |
| `--project ID` | PostHog project id. |
| `--since`, `--until` | The window: any date format ClickHouse parses, such as `2026-09-20`. |
| `--limit N` | Recordings per page. Default 100, at most 1000. |
| `--cursor TOKEN` | Resume after a previous page, using its `next_cursor`. The token carries the project, host and window; flags given with it must match. |
| `--host URL` | Default `https://eu.posthog.com`. |
| `--timeout SECONDS` | Per request. Default 120. |

Each recording in the page has `session_id`, `distinct_id`, `start_time`, `end_time`,
`active_s`, `clicks`, `keypresses`, `console_errors`, `first_url`, `bytes`,
`snapshot_source`, `snapshot_library` and `retention_period_days`.

### recordings fetch

Downloads one recording: `--project ID` and `--session ID`, with `--host`, `--timeout`,
`--max-requests N` (default 50, at most 59) and `--max-wait SECONDS` (default 60; see
[PostHog](#posthog)).

### decode

`--recording FILE` normalizes a recording file into a `recording` artifact. `compile` reads
recording files directly, so you rarely need this.

### compile

| Flag | Meaning |
| --- | --- |
| `--recording FILE` | The recording. |
| `--vocab FILE` | The feature map. |
| `--app ID` | Which app in the map the recording belongs to. |
| `--timings` | Report how long each stage took, on stderr. |

### analyze

Writes a report for a trace. It does one of three things:

| Flag | Does | Network |
| --- | --- | --- |
| `--prepare-only` | Writes the exact model request, as an `analysis_request`. | None |
| `--response FILE` | Checks a saved model answer, and writes an `analysis`. | None |
| `--model ID` | Calls the model through OpenRouter; needs `OPENROUTER_API_KEY`. | OpenRouter |

`--prepare-only` and `--response` take precedence over `--model`. With none of the three, the
command fails: no model is chosen for you.

| Flag | Meaning |
| --- | --- |
| `--trace FILE` | The trace to analyze. |
| `--recording FILE` | Compile this recording first, instead of reading a trace. Needs `--app`. |
| `--vocab FILE` | The feature map. Must be the one the trace was compiled with. |
| `--visit N` | Analyze one visit (counting from 0) instead of the whole trace. |
| `--context FILE` | A header for the model: who the user is, which account. |
| `--system-prompt FILE` | Replace the built-in instructions. The report records the file's name and digest. |
| `--openrouter-url URL`, `--timeout SECONDS` | As for `vocab build`. |

### run

Fetches or reads a recording, compiles it, and analyzes every visit with user actions: `compile`
and `analyze` in one step, writing a `session`.

| Flag | Meaning |
| --- | --- |
| `--recording FILE` | Read this recording. |
| `--session ID`, `--project ID` | Or fetch this PostHog session. Needs `POSTHOG_API_KEY`. |
| `--vocab FILE`, `--app ID` | As for `compile`. |
| `--visit N` | Analyze only this visit. |
| `--save-recording FILE` | Also write the fetched recording, as `recordings fetch` would. |
| `--save-trace FILE` | Also write the trace, as `compile` would. |
| `--prepare-only`, `--model ID`, `--context FILE`, `--system-prompt FILE` | As for `analyze`. |
| `--host`, `--max-requests`, `--max-wait` | As for `recordings fetch`. |

## Environment

| Variable | Needed by |
| --- | --- |
| `POSTHOG_API_KEY` | `recordings list`, `recordings fetch`, `run --session`. A PostHog personal API key with Query Read and Session recording Read access. |
| `OPENROUTER_API_KEY` | Model calls from `vocab build`, `analyze` and `run`. |

No other configuration comes from the environment.

## Output and errors

With `--out FILE`, output is written to a temporary file next to it and then renamed into
place, so a failed write leaves the old file intact. Failures are one JSON object on stderr:

```json
{"error": "reading nope.json: No such file or directory (os error 2)", "retryable": false}
```

| Exit code | Meaning |
| --- | --- |
| `0` | Success. |
| `1` | Invalid input or configuration, or an upstream failure that won't pass. |
| `2` | Invalid invocation: an unknown flag, a missing argument, a bad value. |
| `75` | A transient upstream failure: rate limit (429), server error (5xx), timeout, or a lost connection. `retryable` is `true`; try again later. |

## product.json

```json
{
  "apps": {
    "web": { "project": 12345, "host": "app.example.com", "audience": "workspace admins" }
  }
}
```

One entry per app, keyed by an id you choose and pass as `--app`. `project` is the PostHog
project id, `host` is where the app runs, and `audience` tells the model who uses it. YAML works
too.

## Feature map

`vocab build` writes a snapshot: the map under `vocabulary`, and its provenance, the digest of
the config and of each source file, the source revision, and how it was made (a model's name
and usage, or a validated candidate). `--vocab` also accepts a bare map as YAML or JSON:

```yaml
version: 1
apps:
  workspace: { project: 1, host: app.test, audience: "workspace admins" }
surfaces:
  - { id: workspace.members, app: workspace, route: /settings/members, name: "Members" }
features:
  - id: members.invite.send
    surface: workspace.members
    name: "Send invite"
    matchers: { testid: [send-invite] }
    source: "members.tsx:11"
terms:
  - term: seat
    means: "Every member holds a seat until removed, deactivated members included."
    source: "members.server.ts:4"
statuses:
  - { kind: member, value: deactivated, label: Deactivated }
gaps: ["Billing page: billing.tsx is not among the sources, so seat purchases are unnamed."]
```

| Field | Required | Holds |
| --- | --- | --- |
| `version` | Yes | `1`. |
| `apps` | Yes | As in `product.json`. |
| `surfaces` | Yes | Pages: `id`, `app`, `route`, `name`, and optionally `purpose` and `states` (URL parameters and what they mean). A `route` is a path template where `:name` matches one segment. |
| `features` | | Controls: `id`, `surface`, `name`, `matchers`. |
| `terms` | | The product's words and rules: `term`, `means`, and optionally `ui_labels` and `backend`. |
| `statuses` | | Record states: `kind`, `value`, and optionally `label`. |
| `events` | | Product analytics events: `name`, `app`. |
| `excluded_surfaces` | | Routes left out on purpose: `app`, `route`, `reason`. |
| `gaps` | | What the map doesn't cover, as text. |
| `grid` | | How data tables identify rows; see below. |
| `telemetry` | | Request URL fragments that are telemetry, not product traffic. |
| `error_text` | | Regular expressions for on-screen error messages. |
| `thresholds` | | Timing and distance rules; see below. |

Any entry may carry a `source`: the file and line it came from.

### Matchers

A feature's `matchers` say how to recognize its control. A feature belongs to one surface, a
list of them, or `"*"` for app-wide controls such as a navigation bar, which must also set
`app`. From most to least specific:

| Matcher | Matches |
| --- | --- |
| `testid` | `data-testid` values. |
| `data_attr` | `"attr"` or `"attr=value"`, with or without the `data-` prefix. |
| `aria` | The accessible label. |
| `title` | The `title` attribute. |
| `placeholder` | An input's placeholder. |
| `href` | A link's target. |
| `text` | Visible text. |
| `role` | The ARIA role. |
| `class_contains` | Part of a class name. |
| `aria_template`, `text_template`, `title_template` | Text with `{name}` slots, such as `"Remove {member}"`. |

Unknown matcher keys are rejected.

### Grids

For data tables, `grid` says how to tell rows apart. By default a row is identified by its
label, the first column labels a row, and a cell uses its column's feature when it has none of
its own.

```yaml
grid:
  row_keys: [{ attribute: data-row-id }]    # row identity, tried in order
  label_columns: [[name, title]]            # header words that pick the label column
  generic_cell_features: { prefixes: [grid.column.], suffixes: [-cell] }
```

Without `row_keys`, a re-sorted table and an edited one can look the same.

### Errors and telemetry

Common monitoring requests are ignored, and English error patterns are built in. Add your own:

```yaml
telemetry: [/rum]
error_text: ['(?i)\b(could not|failure)\b']
```

Short text that an action reveals raises the `error_shown` flag when it matches `error_text`,
even when the request behind it succeeded.

### Thresholds

Override any of these; unknown keys are rejected.

| Key | Default | Meaning |
| --- | --- | --- |
| `effect_window_ms` | 2000 | How long after an action its effects are collected. |
| `gesture_ms` | 1000 | A click adopts a mouse-down this recent. |
| `programmatic_input_ms` | 100 | Input on an element this newly shown is the page filling it in, not the user. |
| `thrash_ms` | 10000 | Going A → B → A within this long is `thrash`. |
| `rage_clicks` | 3 | At least this many clicks… |
| `rage_window_ms` | 1000 | …within this long… |
| `rage_px` | 30 | …and this close together is `rage`. |
| `slow_ms` | 1000 | A first visible reaction this late is `slow`. |
| `idle_ms` | 30000 | A pause this long on a visible page is `idle`. |
| `visit_gap_ms` | 1800000 | This long without actions (30 minutes) starts a new visit. |

```yaml
thresholds: { slow_ms: 3000 }
```

## Trace

A trace's `tsv` field holds one line per action:

| Column | Holds |
| --- | --- |
| `ref` | The action's reference: `e1`, `e2`, and so on. Reports cite these. |
| `t_s` | Seconds from the start of the recording. |
| `win` | The browser tab: `w1`, `w2`. Each tab is replayed on its own. |
| `kind` | What happened; see below. |
| `surface` | The feature map surface the page matched. |
| `target` | What was acted on, such as `button "Save"`. In a table, the row and column. |
| `feature` | The feature map feature it matched, if any. |
| `react` | How long the first to the last visible reaction took: `150ms`, `50→150ms`, or `on press`. |
| `effect` | What changed. `none` means nothing visible. |
| `flags` | Signs of trouble; see below. |

The JSON `actions` array holds the same actions in full: timestamps, paths, the resolved
element, click coordinates, structured effects and flags.

| Kind | Meaning |
| --- | --- |
| `click`, `dblclick`, `contextmenu` | A click, double-click, or right-click. |
| `input` | Typing into a field or toggling it. Masked fields reveal only the length. |
| `nav` | The tab moved to another page. |
| `hidden`, `visible` | The tab left view, or came back. |
| `idle` | Nothing happened on a visible page for a while (30 s by default). |
| `console_error` | The page logged an error. |
| `net_error` | A request from the product failed. |

Clicks, double-clicks, right-clicks and inputs are the user's **gestures**.

| Effect | Means |
| --- | --- |
| `text "Save" → "Saved"` | Visible text changed. |
| `cell Status "Alpha": "Queued" → "Scheduled"` | A table cell changed, in row *Alpha*. |
| `-row "Alpha"; +row "Gamma"` | Table rows were removed and added. |
| `+dialog "Delete record?"` | A dialog opened (`-dialog` when it closes). |
| `+status "Saved" (gone after 600ms)` | A message appeared, then went away. |
| `aria-expanded:false→true` | The control changed state. |
| `req /api/notes 200 300ms` | A request, its status and duration. |
| `net 500 /api/save 2400ms` | A failed request. |
| `typed 4 chars (masked)` | Typing into a masked field. |

Flags are set by code, never by a model:

| Flag | When | Threshold |
| --- | --- | --- |
| `dead` | Text that isn't a control was clicked, and nothing visible happened. | |
| `unresponsive` | A control was clicked, and nothing visible happened. | |
| `rage` | Several clicks close together in time and space. | 3 clicks in 1000 ms within 30 px |
| `slow` | The first visible reaction came late. | `slow_ms`, 1000 |
| `error_after` | A console error or failed request during the action. | |
| `error_shown` | The action revealed error text, even if its request succeeded. | `error_text` |
| `thrash` | Going A → B → A. | `thrash_ms`, 10000 |

Clicking into a form field raises neither `dead` nor `unresponsive`, since focusing a field
needn't change the page.

### Visits

A trace splits into **visits** at every 30 minutes without actions, across all tabs. `visits`
lists each one's first and last `ref` and its start and end. `--visit N` picks one; references
stay those of the whole trace.

### Coverage

A trace is only as complete as its `coverage` says:

| Field | Counts |
| --- | --- |
| `events` | Events in the recording, duplicates included. |
| `duplicates` | Events PostHog stored twice, which are dropped. |
| `uninterpreted` | Events read past, by kind, such as mouse moves, scrolls, and unknown rrweb plugins. |
| `malformed` | Events without the shape rrweb gives them, which are skipped. |
| `opaque_mounts` | Iframes, canvases, embeds and shadow roots, whose insides can't be seen. |
| `extensions` | Browser extension content that was left out. Omitted when empty. |
| `unlocated_snapshots` | Page snapshots without a known address. Omitted when zero. |
| `screenshot_only` | Native screens captured only as screenshots. Omitted when false. |

A trace also records the Spoiler version that compiled it, and the SHA-256 of the recording
and of the feature map. `analyze` refuses a trace compiled with a different feature map than
the one it's given.

## Report

A report is an `analysis` artifact. Its `summary` holds the model's answer, completed and
checked by code:

| Field | Holds |
| --- | --- |
| `who` | Who the user seems to be. `--context` informs it. |
| `intent` | What they came to do, with a confidence from 0 to 1. |
| `tasks` | Each attempt at a job; see below. |
| `steps` | Each action: its `refs`, `action`, `response`, `feature` and `outcome`, with timings and changes filled in by code. |
| `friction` | Each problem: its `refs`, `kind`, `what`, `why` and `severity`. |
| `signals` | Filled in by code: every flagged action, with the friction that explains it. |
| `outcome` | Whether the session succeeded (`yes`, `no` or `unknown`), and how. |
| `summary` | A few sentences summing up the session. |

A task is one attempt at a job, across pages, tabs and detours:

| Field | Written by | Holds |
| --- | --- | --- |
| `goal` | Model | The job, such as "Invite priya@example.com to the workspace." |
| `outcome` | Model | `done`, `workaround`, `gave_up` or `unclear`. |
| `obstacle` | Model | What slowed or blocked it, or `null`. |
| `refs` | Model | Every action of the attempt. |
| `start_s`, `end_s`, `active_s` | Code | When it started and ended, and the active time. |
| `path` | Code | The surfaces it passed through. |
| `actions`, `changes` | Code | How many actions it took, and how many data changes it made. |

- Step outcomes are `progressed`, `no_effect`, `error`, `abandoned` or `left`.
- Friction kinds are `dead_click`, `rage_click`, `error`, `slow`, `confusion_loop`,
  `abandonment` and `other`. The first four need the matching flag on the cited actions, or
  they're dropped; the rest are the model's judgement.
- Severity is `blocking`, `degrading` or `cosmetic`. Hypotheses belong in `why`.

### The check

An answer is rejected if it doesn't match the response schema, or if more than 15% of its cited
references don't exist. A live call then gets one retry, with the reason; `--response` fails
without one. An accepted answer keeps its prose, and `check` records the rest:

| Field | Lists |
| --- | --- |
| `bad_refs`, `bad_ref_ratio` | Cited references that don't exist in the trace. |
| `dropped_friction` | Friction claims whose cited actions lack the matching flag. |
| `unknown_features` | Feature ids the feature map doesn't define. |
| `corrected` | Repairs, such as a step dropped for citing no valid references. |
| `unverified_numbers` | Durations the model stated that the trace doesn't support. |
| `unexplained_signals` | Flagged actions no friction item cites. |
| `uncited_gestures` | User gestures no step cites. |

A report also records its inputs: `trace_digest`, `vocab_digest`, the instructions' name and
SHA-256 (`prompt`), and a `request_digest` over the exact request. A live call adds `model`: the
model that answered, the attempts, `tokens_in`, `tokens_out` and `cost_usd`.

## Artifacts

Every artifact has a `kind` and a `schema_version`. Readers reject unknown kinds and versions.

| Kind | Version | Written by |
| --- | --- | --- |
| `vocabulary_snapshot` | 1 | `vocab build` |
| `vocabulary_check` | 1 | `vocab check` |
| `recording_page` | 2 | `recordings list` |
| `recording` | 1 | `recordings fetch`, `decode` |
| `trace` | 3 | `compile` |
| `analysis_request` | 2 | `analyze --prepare-only` |
| `analysis` | 2 | `analyze --response`, `analyze --model` |
| `session` | 1 | `run` |

A `session` holds the recording's `source`, its `trace`, and `visits`: for each visit analyzed,
its index and either an `analysis`, or a `request` with `--prepare-only`.

## PostHog

- Recordings are read from `https://eu.posthog.com` unless you pass `--host`, for US or
  self-hosted PostHog.
- `recordings list` shows only recordings that started more than 24 hours ago and ended by
  `--until`. Choose windows that have settled, and overlap consecutive ones when syncing.
- Parts of a recording more than 7 days outside the window can make it look complete when it
  isn't.
- Listings and fetches share PostHog's rate limit for a key (paid plans: 60 a minute, 300 an
  hour). A paid key fetches about 75 typical recordings an hour.
- On a rate limit, `recordings fetch` and `run --session` wait as long as PostHog asks, up to
  `--max-wait` seconds in total (default 60; `0` never waits), then exit with `75`.
  `recordings list` doesn't retry.

## Limits and security

- Recordings are untrusted input. `--max-input-mib` bounds what a recording decodes to,
  including compressed fields, and responses from PostHog's query API and from models are
  capped at 32 MiB each.
- Credentials travel only over HTTPS, and redirects are never followed.
- Native iOS and Android taps are matched against PostHog's wireframes, and native routes are
  screen names such as `SettingsScreen`. Screens captured only as screenshots, including
  Flutter and React Native, can't be read: a tap on one appears as `screen (x,y)`.
- Report security issues privately through
  [GitHub security advisories](https://github.com/sahil-shubham/spoiler/security/advisories/new).
