Start

Introduction

The four things Spoiler works with, what each step does with them, and what leaves your machine.

Spoiler is a command-line tool. It reads session recordings of your product and writes, for each one, a report of what the person was trying to do, how it ended, and what got in the way.

It works with four things: a recording, a feature map, a trace and a report. This page explains each one, starting from the recording. Set up puts them to work in your project.

your source code         ─▶  vocab build  ─▶  feature map
recording + feature map  ─▶  compile      ─▶  trace
trace + feature map      ─▶  analyze      ─▶  report

Recordings

A session recording is a log of one person’s visit to your product. It isn’t a video. The recorder writes down the page itself, as HTML, and then every change to it, along with each click, each field typed into, and, when enabled, network requests and console errors. Because the log holds the page’s actual elements, Spoiler can tell exactly which button was clicked and what changed on screen afterwards.

The format is rrweb, the open-source recorder that PostHog’s session replay is built on. Spoiler fetches recordings from PostHog directly, from web apps and from native iOS and Android apps, or reads rrweb files you export yourself.

The feature map

A recording knows that someone clicked <button data-testid="send-invite">. It doesn’t know that this is Send invite, or that invites fail while every seat is taken. The feature map holds that knowledge about your product:

  • Surfaces: your pages, by route, such as /settings/members.
  • Features: the controls on them, and how to recognize each one in a recording, such as “the element with data-testid="send-invite"”.
  • Terms: your product’s rules and vocabulary, such as what a seat is and when it’s freed.
  • Gaps: what the map doesn’t cover, listed rather than guessed.

You make it once, with spoiler vocab build. You name the source files that define your routes, controls and rules, and a model drafts the map from them. Every entry cites the file and line it came from. You review the map, correct it if needed, and reuse it for every recording until those files change. Each trace records exactly which map it was built with.

The CLI calls the feature map a vocabulary: its commands are spoiler vocab build and spoiler vocab check.

The trace

spoiler compile turns a recording into a trace: one line per action.

ref  t_s  win  kind   surface    target         feature  react  effect                  flags
e1   0.0  w1   nav    demo.page  /page
e2   1.1  w1   click  demo.page  button "Save"           150ms  text "Save" → "Saved"

Each line gives the action a reference (e2) and says when it happened, what was acted on, which feature that was, how quickly the page reacted, and what changed. Here, someone clicked Save, and 150 ms later the button read Saved.

The last column holds flags: signs of trouble that code detects, such as a click that changed nothing, rapid repeated clicks, a slow reaction, or an error appearing right after an action. A model never adds or removes flags.

Compiling is plain code. It needs no model and no network, and the same recording, feature map and Spoiler version always give the same trace.

The report

spoiler analyze sends the trace, and the parts of the feature map it touched, to a model you choose. The recording itself is never sent. The model groups the actions into:

  • Tasks: what the person set out to do, how it ended (done, workaround, gave_up or unclear), and what stood in the way.
  • Steps: each action and how the product responded.
  • Friction: each problem, its kind and severity, and a hypothesis for why it happened.

Every item cites the trace references behind it. Then code checks the answer:

  • An answer that doesn’t match the expected shape, or where more than 15% of the cited references don’t exist, is rejected. A live model call gets one retry with the reason.
  • A claim of a dead click, rage click, error or slow response is dropped unless the actions it cites carry the matching flag.
  • Timings, paths and counts are filled in from the trace, never taken from the model.
  • Anything else that couldn’t be verified is listed in the report’s check.

So the measurements in a report come from code, and the explanations come from a model and are marked as such. spoiler run does both steps for one recording: it compiles it and writes a report for each visit with user actions.

What leaves your machine

Step Sends To
vocab build The source files you name OpenRouter, to the model you choose
recordings list, recordings fetch Read requests PostHog
compile Nothing
analyze, run with --model The trace and the feature map entries it touched OpenRouter, to the model you choose

Spoiler never picks a model for you, and model calls ask OpenRouter for zero-data-retention routing. Keys are read only from the environment. Recordings contain what users saw and typed: review what you send, and don’t commit recordings to your repository.

What Spoiler can’t see

Spoiler sees only what the recording captured:

  • Content inside iframes, drawings on a canvas, and the insides of web components that use a shadow root.
  • Key presses. rrweb doesn’t record them, so an action taken with a keyboard shortcut shows up only through what it changed.
  • The contents of masked inputs, such as passwords. Spoiler sees only their length.

Each trace counts what it couldn’t see in its coverage; see the trace reference.