Start
Introduction
The four things Spoiler works with, what each step does with them, and what leaves your machine.
Spoiler is a command-line tool. It reads session recordings of your product and writes, for each one, a report of what the person was trying to do, how it ended, and what got in the way.
It works with four things: a recording, a feature map, a trace and a report. This page explains each one, starting from the recording. Set up puts them to work in your project.
your source code ─▶ vocab build ─▶ feature map
recording + feature map ─▶ compile ─▶ trace
trace + feature map ─▶ analyze ─▶ report
Recordings
A session recording is a log of one person’s visit to your product. It isn’t a video. The recorder writes down the page itself, as HTML, and then every change to it, along with each click, each field typed into, and, when enabled, network requests and console errors. Because the log holds the page’s actual elements, Spoiler can tell exactly which button was clicked and what changed on screen afterwards.
The format is rrweb, the open-source recorder that PostHog’s session replay is built on. Spoiler fetches recordings from PostHog directly, from web apps and from native iOS and Android apps, or reads rrweb files you export yourself.
The feature map
A recording knows that someone clicked <button data-testid="send-invite">. It doesn’t know
that this is Send invite, or that invites fail while every seat is taken. The feature map
holds that knowledge about your product:
- Surfaces: your pages, by route, such as
/settings/members. - Features: the controls on them, and how to recognize each one in a recording, such as
“the element with
data-testid="send-invite"”. - Terms: your product’s rules and vocabulary, such as what a seat is and when it’s freed.
- Gaps: what the map doesn’t cover, listed rather than guessed.
You make it once, with spoiler vocab build. You name the source files that define your
routes, controls and rules, and a model drafts the map from them. Every entry cites the file
and line it came from. You review the map, correct it if needed, and reuse it for every
recording until those files change. Each trace records exactly which map it was built with.
The CLI calls the feature map a vocabulary: its commands are spoiler vocab build and
spoiler vocab check.
The trace
spoiler compile turns a recording into a trace: one line per action.
ref t_s win kind surface target feature react effect flags
e1 0.0 w1 nav demo.page /page
e2 1.1 w1 click demo.page button "Save" 150ms text "Save" → "Saved"
Each line gives the action a reference (e2) and says when it happened, what was acted on,
which feature that was, how quickly the page reacted, and what changed. Here, someone clicked
Save, and 150 ms later the button read Saved.
The last column holds flags: signs of trouble that code detects, such as a click that changed nothing, rapid repeated clicks, a slow reaction, or an error appearing right after an action. A model never adds or removes flags.
Compiling is plain code. It needs no model and no network, and the same recording, feature map and Spoiler version always give the same trace.
The report
spoiler analyze sends the trace, and the parts of the feature map it touched, to a model you
choose. The recording itself is never sent. The model groups the actions into:
- Tasks: what the person set out to do, how it ended (
done,workaround,gave_uporunclear), and what stood in the way. - Steps: each action and how the product responded.
- Friction: each problem, its kind and severity, and a hypothesis for why it happened.
Every item cites the trace references behind it. Then code checks the answer:
- An answer that doesn’t match the expected shape, or where more than 15% of the cited references don’t exist, is rejected. A live model call gets one retry with the reason.
- A claim of a dead click, rage click, error or slow response is dropped unless the actions it cites carry the matching flag.
- Timings, paths and counts are filled in from the trace, never taken from the model.
- Anything else that couldn’t be verified is listed in the report’s
check.
So the measurements in a report come from code, and the explanations come from a model and are
marked as such. spoiler run does both steps for one recording: it compiles it and writes a
report for each visit with user actions.
What leaves your machine
| Step | Sends | To |
|---|---|---|
vocab build |
The source files you name | OpenRouter, to the model you choose |
recordings list, recordings fetch |
Read requests | PostHog |
compile |
Nothing | |
analyze, run with --model |
The trace and the feature map entries it touched | OpenRouter, to the model you choose |
Spoiler never picks a model for you, and model calls ask OpenRouter for zero-data-retention routing. Keys are read only from the environment. Recordings contain what users saw and typed: review what you send, and don’t commit recordings to your repository.
What Spoiler can’t see
Spoiler sees only what the recording captured:
- Content inside iframes, drawings on a canvas, and the insides of web components that use a shadow root.
- Key presses. rrweb doesn’t record them, so an action taken with a keyboard shortcut shows up only through what it changed.
- The contents of masked inputs, such as passwords. Spoiler sees only their length.
Each trace counts what it couldn’t see in its coverage; see
the trace reference.