Reference
Reference
Every command and flag, file format, and output field.
Everything Spoiler reads and writes, in one place. Introduction explains what each part
is for; Set up puts them together. spoiler <command> --help always lists the
flags of the version you have installed.
Installing
curl -fsSL https://spoiler.sh/install | sh
The script picks the build for your OS and CPU, checks its SHA-256, and installs it to
~/.local/bin. Linux gets a static musl build that runs on any distribution. Pin a release with
curl -fsSL https://spoiler.sh/install | SPOILER_VERSION=v0.1.0 sh, and choose the directory
with SPOILER_INSTALL_DIR.
| Installer | Command |
|---|---|
| uv | uv tool install spoiler |
| pip | pip install spoiler |
| Cargo (Rust 1.88+) | cargo install spoiler --locked |
The Python packages only put the spoiler binary on PATH; there’s no Python API.
GitHub Releases has an archive for each
target, each with a .sha256 file:
aarch64-apple-darwin,x86_64-apple-darwinx86_64-unknown-linux-gnu,aarch64-unknown-linux-gnu(glibc 2.28 or later)x86_64-unknown-linux-musl,aarch64-unknown-linux-musl(static; any Linux)
From a checkout of the repository: cargo install --path crates/cli --locked.
Commands
| Command | Reads | Writes | Network |
|---|---|---|---|
vocab build |
product.json and source files |
vocabulary_snapshot |
OpenRouter, unless --candidate |
vocab check |
A feature map | vocabulary_check |
None |
recordings list |
A PostHog project | recording_page |
PostHog |
recordings fetch |
A PostHog session | recording |
PostHog |
decode |
A recording file | recording |
None |
compile |
A recording and a feature map | trace |
None |
analyze |
A trace, or a recording, and a feature map | analysis_request or analysis |
OpenRouter with --model |
run |
A recording or PostHog session, and a feature map | session |
PostHog with --session, OpenRouter with --model |
Every command:
- writes one JSON artifact, to
--out FILEor to stdout; - reads
-as standard input for any input path, so commands pipe into each other; - takes all configuration as flags, and credentials only from the environment;
- accepts
--max-input-mib(default 512), the most a recording may decode to.
spoiler recordings fetch --project 12345 --session "$SESSION" \
| spoiler compile --recording - --vocab spoiler/vocab.json --app web \
| spoiler analyze --trace - --vocab spoiler/vocab.json --prepare-only
vocab build
Builds a feature map from product.json and the source files you name.
| Flag | Meaning |
|---|---|
--config FILE |
The product config. Required. |
--source FILE |
A source file for the model to read. Required; repeat for each file. |
--model ID |
OpenRouter model to draft the map. Needs OPENROUTER_API_KEY. |
--candidate FILE |
Package this feature map instead of drafting one; no model call. Takes precedence over --model. |
--source-revision REV |
The revision the sources were read at, such as a commit, recorded with the map. |
--openrouter-url URL |
An OpenRouter-compatible endpoint. Default https://openrouter.ai/api/v1. |
--timeout SECONDS |
Per request. Default 120. |
vocab check
--vocab FILE validates a feature map, in YAML, JSON, or a built snapshot, and reports its
digest, apps, counts, and warnings for entries that can never match. No network.
recordings list
One page of recordings that started in a time window, oldest first.
| Flag | Meaning |
|---|---|
--project ID |
PostHog project id. |
--since, --until |
The window: any date format ClickHouse parses, such as 2026-09-20. |
--limit N |
Recordings per page. Default 100, at most 1000. |
--cursor TOKEN |
Resume after a previous page, using its next_cursor. The token carries the project, host and window; flags given with it must match. |
--host URL |
Default https://eu.posthog.com. |
--timeout SECONDS |
Per request. Default 120. |
Each recording in the page has session_id, distinct_id, start_time, end_time,
active_s, clicks, keypresses, console_errors, first_url, bytes,
snapshot_source, snapshot_library and retention_period_days.
recordings fetch
Downloads one recording: --project ID and --session ID, with --host, --timeout,
--max-requests N (default 50, at most 59) and --max-wait SECONDS (default 60; see
PostHog).
decode
--recording FILE normalizes a recording file into a recording artifact. compile reads
recording files directly, so you rarely need this.
compile
| Flag | Meaning |
|---|---|
--recording FILE |
The recording. |
--vocab FILE |
The feature map. |
--app ID |
Which app in the map the recording belongs to. |
--timings |
Report how long each stage took, on stderr. |
analyze
Writes a report for a trace. It does one of three things:
| Flag | Does | Network |
|---|---|---|
--prepare-only |
Writes the exact model request, as an analysis_request. |
None |
--response FILE |
Checks a saved model answer, and writes an analysis. |
None |
--model ID |
Calls the model through OpenRouter; needs OPENROUTER_API_KEY. |
OpenRouter |
--prepare-only and --response take precedence over --model. With none of the three, the
command fails: no model is chosen for you.
| Flag | Meaning |
|---|---|
--trace FILE |
The trace to analyze. |
--recording FILE |
Compile this recording first, instead of reading a trace. Needs --app. |
--vocab FILE |
The feature map. Must be the one the trace was compiled with. |
--visit N |
Analyze one visit (counting from 0) instead of the whole trace. |
--context FILE |
A header for the model: who the user is, which account. |
--system-prompt FILE |
Replace the built-in instructions. The report records the file’s name and digest. |
--openrouter-url URL, --timeout SECONDS |
As for vocab build. |
run
Fetches or reads a recording, compiles it, and analyzes every visit with user actions: compile
and analyze in one step, writing a session.
| Flag | Meaning |
|---|---|
--recording FILE |
Read this recording. |
--session ID, --project ID |
Or fetch this PostHog session. Needs POSTHOG_API_KEY. |
--vocab FILE, --app ID |
As for compile. |
--visit N |
Analyze only this visit. |
--save-recording FILE |
Also write the fetched recording, as recordings fetch would. |
--save-trace FILE |
Also write the trace, as compile would. |
--prepare-only, --model ID, --context FILE, --system-prompt FILE |
As for analyze. |
--host, --max-requests, --max-wait |
As for recordings fetch. |
Environment
| Variable | Needed by |
|---|---|
POSTHOG_API_KEY |
recordings list, recordings fetch, run --session. A PostHog personal API key with Query Read and Session recording Read access. |
OPENROUTER_API_KEY |
Model calls from vocab build, analyze and run. |
No other configuration comes from the environment.
Output and errors
With --out FILE, output is written to a temporary file next to it and then renamed into
place, so a failed write leaves the old file intact. Failures are one JSON object on stderr:
{"error": "reading nope.json: No such file or directory (os error 2)", "retryable": false}
| Exit code | Meaning |
|---|---|
0 |
Success. |
1 |
Invalid input or configuration, or an upstream failure that won’t pass. |
2 |
Invalid invocation: an unknown flag, a missing argument, a bad value. |
75 |
A transient upstream failure: rate limit (429), server error (5xx), timeout, or a lost connection. retryable is true; try again later. |
product.json
{
"apps": {
"web": { "project": 12345, "host": "app.example.com", "audience": "workspace admins" }
}
}
One entry per app, keyed by an id you choose and pass as --app. project is the PostHog
project id, host is where the app runs, and audience tells the model who uses it. YAML works
too.
Feature map
vocab build writes a snapshot: the map under vocabulary, and its provenance, the digest of
the config and of each source file, the source revision, and how it was made (a model’s name
and usage, or a validated candidate). --vocab also accepts a bare map as YAML or JSON:
version: 1
apps:
workspace: { project: 1, host: app.test, audience: "workspace admins" }
surfaces:
- { id: workspace.members, app: workspace, route: /settings/members, name: "Members" }
features:
- id: members.invite.send
surface: workspace.members
name: "Send invite"
matchers: { testid: [send-invite] }
source: "members.tsx:11"
terms:
- term: seat
means: "Every member holds a seat until removed, deactivated members included."
source: "members.server.ts:4"
statuses:
- { kind: member, value: deactivated, label: Deactivated }
gaps: ["Billing page: billing.tsx is not among the sources, so seat purchases are unnamed."]
| Field | Required | Holds |
|---|---|---|
version |
Yes | 1. |
apps |
Yes | As in product.json. |
surfaces |
Yes | Pages: id, app, route, name, and optionally purpose and states (URL parameters and what they mean). A route is a path template where :name matches one segment. |
features |
Controls: id, surface, name, matchers. |
|
terms |
The product’s words and rules: term, means, and optionally ui_labels and backend. |
|
statuses |
Record states: kind, value, and optionally label. |
|
events |
Product analytics events: name, app. |
|
excluded_surfaces |
Routes left out on purpose: app, route, reason. |
|
gaps |
What the map doesn’t cover, as text. | |
grid |
How data tables identify rows; see below. | |
telemetry |
Request URL fragments that are telemetry, not product traffic. | |
error_text |
Regular expressions for on-screen error messages. | |
thresholds |
Timing and distance rules; see below. |
Any entry may carry a source: the file and line it came from.
Matchers
A feature’s matchers say how to recognize its control. A feature belongs to one surface, a
list of them, or "*" for app-wide controls such as a navigation bar, which must also set
app. From most to least specific:
| Matcher | Matches |
|---|---|
testid |
data-testid values. |
data_attr |
"attr" or "attr=value", with or without the data- prefix. |
aria |
The accessible label. |
title |
The title attribute. |
placeholder |
An input’s placeholder. |
href |
A link’s target. |
text |
Visible text. |
role |
The ARIA role. |
class_contains |
Part of a class name. |
aria_template, text_template, title_template |
Text with {name} slots, such as "Remove {member}". |
Unknown matcher keys are rejected.
Grids
For data tables, grid says how to tell rows apart. By default a row is identified by its
label, the first column labels a row, and a cell uses its column’s feature when it has none of
its own.
grid:
row_keys: [{ attribute: data-row-id }] # row identity, tried in order
label_columns: [[name, title]] # header words that pick the label column
generic_cell_features: { prefixes: [grid.column.], suffixes: [-cell] }
Without row_keys, a re-sorted table and an edited one can look the same.
Errors and telemetry
Common monitoring requests are ignored, and English error patterns are built in. Add your own:
telemetry: [/rum]
error_text: ['(?i)\b(could not|failure)\b']
Short text that an action reveals raises the error_shown flag when it matches error_text,
even when the request behind it succeeded.
Thresholds
Override any of these; unknown keys are rejected.
| Key | Default | Meaning |
|---|---|---|
effect_window_ms |
2000 | How long after an action its effects are collected. |
gesture_ms |
1000 | A click adopts a mouse-down this recent. |
programmatic_input_ms |
100 | Input on an element this newly shown is the page filling it in, not the user. |
thrash_ms |
10000 | Going A → B → A within this long is thrash. |
rage_clicks |
3 | At least this many clicks… |
rage_window_ms |
1000 | …within this long… |
rage_px |
30 | …and this close together is rage. |
slow_ms |
1000 | A first visible reaction this late is slow. |
idle_ms |
30000 | A pause this long on a visible page is idle. |
visit_gap_ms |
1800000 | This long without actions (30 minutes) starts a new visit. |
thresholds: { slow_ms: 3000 }
Trace
A trace’s tsv field holds one line per action:
| Column | Holds |
|---|---|
ref |
The action’s reference: e1, e2, and so on. Reports cite these. |
t_s |
Seconds from the start of the recording. |
win |
The browser tab: w1, w2. Each tab is replayed on its own. |
kind |
What happened; see below. |
surface |
The feature map surface the page matched. |
target |
What was acted on, such as button "Save". In a table, the row and column. |
feature |
The feature map feature it matched, if any. |
react |
How long the first to the last visible reaction took: 150ms, 50→150ms, or on press. |
effect |
What changed. none means nothing visible. |
flags |
Signs of trouble; see below. |
The JSON actions array holds the same actions in full: timestamps, paths, the resolved
element, click coordinates, structured effects and flags.
| Kind | Meaning |
|---|---|
click, dblclick, contextmenu |
A click, double-click, or right-click. |
input |
Typing into a field or toggling it. Masked fields reveal only the length. |
nav |
The tab moved to another page. |
hidden, visible |
The tab left view, or came back. |
idle |
Nothing happened on a visible page for a while (30 s by default). |
console_error |
The page logged an error. |
net_error |
A request from the product failed. |
Clicks, double-clicks, right-clicks and inputs are the user’s gestures.
| Effect | Means |
|---|---|
text "Save" → "Saved" |
Visible text changed. |
cell Status "Alpha": "Queued" → "Scheduled" |
A table cell changed, in row Alpha. |
-row "Alpha"; +row "Gamma" |
Table rows were removed and added. |
+dialog "Delete record?" |
A dialog opened (-dialog when it closes). |
+status "Saved" (gone after 600ms) |
A message appeared, then went away. |
aria-expanded:false→true |
The control changed state. |
req /api/notes 200 300ms |
A request, its status and duration. |
net 500 /api/save 2400ms |
A failed request. |
typed 4 chars (masked) |
Typing into a masked field. |
Flags are set by code, never by a model:
| Flag | When | Threshold |
|---|---|---|
dead |
Text that isn’t a control was clicked, and nothing visible happened. | |
unresponsive |
A control was clicked, and nothing visible happened. | |
rage |
Several clicks close together in time and space. | 3 clicks in 1000 ms within 30 px |
slow |
The first visible reaction came late. | slow_ms, 1000 |
error_after |
A console error or failed request during the action. | |
error_shown |
The action revealed error text, even if its request succeeded. | error_text |
thrash |
Going A → B → A. | thrash_ms, 10000 |
Clicking into a form field raises neither dead nor unresponsive, since focusing a field
needn’t change the page.
Visits
A trace splits into visits at every 30 minutes without actions, across all tabs. visits
lists each one’s first and last ref and its start and end. --visit N picks one; references
stay those of the whole trace.
Coverage
A trace is only as complete as its coverage says:
| Field | Counts |
|---|---|
events |
Events in the recording, duplicates included. |
duplicates |
Events PostHog stored twice, which are dropped. |
uninterpreted |
Events read past, by kind, such as mouse moves, scrolls, and unknown rrweb plugins. |
malformed |
Events without the shape rrweb gives them, which are skipped. |
opaque_mounts |
Iframes, canvases, embeds and shadow roots, whose insides can’t be seen. |
extensions |
Browser extension content that was left out. Omitted when empty. |
unlocated_snapshots |
Page snapshots without a known address. Omitted when zero. |
screenshot_only |
Native screens captured only as screenshots. Omitted when false. |
A trace also records the Spoiler version that compiled it, and the SHA-256 of the recording
and of the feature map. analyze refuses a trace compiled with a different feature map than
the one it’s given.
Report
A report is an analysis artifact. Its summary holds the model’s answer, completed and
checked by code:
| Field | Holds |
|---|---|
who |
Who the user seems to be. --context informs it. |
intent |
What they came to do, with a confidence from 0 to 1. |
tasks |
Each attempt at a job; see below. |
steps |
Each action: its refs, action, response, feature and outcome, with timings and changes filled in by code. |
friction |
Each problem: its refs, kind, what, why and severity. |
signals |
Filled in by code: every flagged action, with the friction that explains it. |
outcome |
Whether the session succeeded (yes, no or unknown), and how. |
summary |
A few sentences summing up the session. |
A task is one attempt at a job, across pages, tabs and detours:
| Field | Written by | Holds |
|---|---|---|
goal |
Model | The job, such as “Invite priya@example.com to the workspace.” |
outcome |
Model | done, workaround, gave_up or unclear. |
obstacle |
Model | What slowed or blocked it, or null. |
refs |
Model | Every action of the attempt. |
start_s, end_s, active_s |
Code | When it started and ended, and the active time. |
path |
Code | The surfaces it passed through. |
actions, changes |
Code | How many actions it took, and how many data changes it made. |
- Step outcomes are
progressed,no_effect,error,abandonedorleft. - Friction kinds are
dead_click,rage_click,error,slow,confusion_loop,abandonmentandother. The first four need the matching flag on the cited actions, or they’re dropped; the rest are the model’s judgement. - Severity is
blocking,degradingorcosmetic. Hypotheses belong inwhy.
The check
An answer is rejected if it doesn’t match the response schema, or if more than 15% of its cited
references don’t exist. A live call then gets one retry, with the reason; --response fails
without one. An accepted answer keeps its prose, and check records the rest:
| Field | Lists |
|---|---|
bad_refs, bad_ref_ratio |
Cited references that don’t exist in the trace. |
dropped_friction |
Friction claims whose cited actions lack the matching flag. |
unknown_features |
Feature ids the feature map doesn’t define. |
corrected |
Repairs, such as a step dropped for citing no valid references. |
unverified_numbers |
Durations the model stated that the trace doesn’t support. |
unexplained_signals |
Flagged actions no friction item cites. |
uncited_gestures |
User gestures no step cites. |
A report also records its inputs: trace_digest, vocab_digest, the instructions’ name and
SHA-256 (prompt), and a request_digest over the exact request. A live call adds model: the
model that answered, the attempts, tokens_in, tokens_out and cost_usd.
Artifacts
Every artifact has a kind and a schema_version. Readers reject unknown kinds and versions.
| Kind | Version | Written by |
|---|---|---|
vocabulary_snapshot |
1 | vocab build |
vocabulary_check |
1 | vocab check |
recording_page |
2 | recordings list |
recording |
1 | recordings fetch, decode |
trace |
3 | compile |
analysis_request |
2 | analyze --prepare-only |
analysis |
2 | analyze --response, analyze --model |
session |
1 | run |
A session holds the recording’s source, its trace, and visits: for each visit analyzed,
its index and either an analysis, or a request with --prepare-only.
PostHog
- Recordings are read from
https://eu.posthog.comunless you pass--host, for US or self-hosted PostHog. recordings listshows only recordings that started more than 24 hours ago and ended by--until. Choose windows that have settled, and overlap consecutive ones when syncing.- Parts of a recording more than 7 days outside the window can make it look complete when it isn’t.
- Listings and fetches share PostHog’s rate limit for a key (paid plans: 60 a minute, 300 an hour). A paid key fetches about 75 typical recordings an hour.
- On a rate limit,
recordings fetchandrun --sessionwait as long as PostHog asks, up to--max-waitseconds in total (default 60;0never waits), then exit with75.recordings listdoesn’t retry.
Limits and security
- Recordings are untrusted input.
--max-input-mibbounds what a recording decodes to, including compressed fields, and responses from PostHog’s query API and from models are capped at 32 MiB each. - Credentials travel only over HTTPS, and redirects are never followed.
- Native iOS and Android taps are matched against PostHog’s wireframes, and native routes are
screen names such as
SettingsScreen. Screens captured only as screenshots, including Flutter and React Native, can’t be read: a tap on one appears asscreen (x,y). - Report security issues privately through GitHub security advisories.