Skip to main content
These are the Reticle tools that do something other than look, act, observe and assert: control coverage, client storage, a fake clock, autonomous crawling, semantic and visual baselines, network mocking, flow recording and replay, and the CI artifact. Reach for them when the verify loop is not the shape of the question. They are not advertised on the default tool surface, but they are reachable from it: reticle_run({ tool, args }) dispatches to any registered tool by name, advertised or not, which is the spelling used on this page. To have them advertised outright instead, with their output schemas, start the daemon with RETICLE_ADVERTISE_ALL_TOOLS=1 (read once at startup, so restart it) and call them by name. Most of the time you want the verify loop. Look, act, observe, assert. These are the tools for the times you want something else, and they are the ones people are most surprised exist.

Available anywhere

These work through the always-on SDK, in any connected session.

Coverage: what you have not driven yet

That is a real capture with 39 of the 41 untouched entries trimmed out. Two controls driven out of forty-three. Not code coverage. Control coverage: which interactive elements this session has actually touched. Useful at the end of a drive, when the agent is about to report success. “I verified the page” reads differently next to exercised: 0.

Storage: the persistence layer

Note the token redaction. Credential-shaped values are stripped here exactly as they are in reticle_network, so reading storage never puts a session token in your context. localStorage, sessionStorage and readable cookies. This is where the “logged in but nothing persisted” bug lives. A login can fire its success signal, mutate state, and write nothing. Reload, and you are back at the login screen.
storageKeysChanged already appears in every act_and_wait summary, so you often do not need this tool. Read it there first.

Clock: freeze time, skip the wait

advanceMs answers { "frozen": true }, because advancing keeps the clock frozen at the new time. reset answers { "frozen": false } and hands time back to the browser. Toasts that auto-dismiss, debounced search, polling, session timeouts, retry backoff. All of these are normally verified by sleeping, which is slow and flaky in equal measure. Freeze the clock, advance it by exactly the interval, assert the consequence. No sleeping, no guessing, and the same result on a fast laptop and a loaded CI runner.

Crawl: click everything, report anomalies

Read coverageNote before you read the counts. interactiveFound is measured when the crawl starts, so on any app where clicking reveals new controls it is a floor rather than a total, and the tool says so rather than letting 43 read as complete. Autonomously clicks every reachable control and reports what went wrong. No script, no plan. The counts are the useful part: deadControls and contradictions are the two that indicate a real problem rather than a busy page. This is the tool for an app you did not write and do not understand yet.
It clicks everything, so point it at a dev environment. maxSteps bounds it, defaulting to 25.

Reconcile: what the API said versus what rendered

Compares the data an endpoint returned against what the page actually displays. compared is the field that makes the empty mismatches meaningful: compared: 0 would mean it had nothing to check, not that everything agreed. The bug it catches is specific and common: the API returned ten rows, the table shows nine, and nothing errored. Neither the network log nor the DOM is wrong on its own. Only the comparison is.

Domain: which of your flows actually prove anything

The most useful tool nobody knows about.
Read that summary again. Thirty-one of forty-seven flows assert no consequence, which means they replay green whatever the app does. That is a suite that looks like coverage and is not. Each flow comes back graded, with the consequence that must hold and a risk level:
declaredUntestedSignals is the other half: signals your app emits that no flow ever checks. It is a to-do list for your test suite, generated from what the app says about itself. The full response also carries declared (the app’s registered testids, signals and stores), riskRanked (every flow ordered worst first) and unassertedFlows.
Run this before writing a new flow. It tells you what is already covered, what is covered badly, and which declared signal has never been asserted. All three are better starting points than guessing.

Baselines: semantic snapshots, not pixels

Records the meaningful state of a page and compares later. Structure and content rather than pixels, so a font change does not fail your check and a missing row does. action: "list" returns the saved names.

Needs a driven browser

Four tools apply their effects through the Chrome DevTools Protocol. The always-on SDK cannot do that, so they need a browser Reticle is driving.
A pooled lease is not enough. Re-verified on 2026-08-16: acquire one with reticle_lease and reticle_screenshot answers
while reticle_network_mock answers { "applied": false, "count": 0, "ok": false, "reason": "no-cdp-provider" }. Under a driven browser, all four work: the captures below were taken that way.Run reticle drive <url>, or point RETICLE_CDP_URL at a Chrome started with --remote-debugging-port.

Screenshots and visual diffing

dimensionMismatch is the field that saves an afternoon. A diff between two different viewport sizes is meaningless, and this says so rather than reporting a huge ratio.
Note the asymmetry. reticle_screenshot names the baseline with name; reticle_visual_diff refers to it with baseline. Passing name to the diff is rejected, and it is an easy mistake to make twice.
fullPage captures the whole scroll height. ref or clip scopes to one element or region. threshold sets the pixel-difference tolerance, defaulting to 0.01, maxRatio caps the acceptable changed fraction, and masks excludes regions that are expected to change. Reticle reads program truth, not pixels, so this is the deliberate exception: a font that failed to load or a compositing glitch is only visible in the actual frame.

Viewport pinning

Fixes the viewport so a visual baseline is reproducible across machines. Without it, a diff taken on a laptop and re-run in CI compares two different layouts and fails for a reason nobody wants to debug.

Network mocking

Return a 500, force offline, or delay a response. First matching rule wins. clear: true turns it off and answers { "applied": true, "count": 0 }. Most error states have never actually run. This is how you find out whether yours works, without touching the backend or waiting for a real outage to tell you.

Flows: record once, replay forever

The loop closer, and the reason Reticle is not only an interactive tool.
The argument names are not interchangeable. reticle_record takes recordingName, reticle_flow_save takes flowName, and neither accepts name. Reticle rejects the call rather than guessing, which is the right call and still costs you a turn.
Stopping a recording used to return a very large payload, almost all of it the raw events array. It no longer does. The timeline is omitted and replaced with a pointer, so a stop is under a kilobyte. This is a real capture of a one-click recording, complete rather than trimmed:
Each step’s args is elided above because a recorded step is an anchor, not a call you make by hand: it resolves a testid at replay time rather than a ref, so copying it into reticle_act would not work. Nine hundred and fifty-three bytes for 445 recorded events. timeline_omitted hands you the exact reticle_observe call to make if you want them, so nothing is lost, it is only not charged for by default.
reticle_record takes action, recordingName and sessionId and nothing else. max_events and filters are rejected, which no longer matters now that the timeline is omitted anyway.
proposedConsequences is the useful part and the reason to record at all. Reticle watched the interaction and worked out what would have proved it, ranked by strength. Tier 0 is the app’s own signal. You do not have to invent the assertion; you pick one.

Saving and replaying

Reticle grades the flow as you save it, and tells you when you have just created a test that cannot fail. Take the warning seriously: an assertion-free flow replays green whatever the app does. A replay against a drifted app names what moved and where. This is reticle_verify {action:"flows"} reporting the same flow after the app changed underneath it:
Passing flows are counted, not described. Only failures carry detail, because the detail is the actionable part. flaky lists flows seen to both pass and fail on unchanged code, which is a bug in the flow rather than in the app. reticle_verify { action: "heal" } proposes the rebind. It refuses when it is not confident:
Flows live in .reticle/ as human-readable files, so they are reviewed in pull requests and diffed like code. They anchor on meaning, a testid plus a signal, rather than volatile refs or coordinates.

Flows in full

Recording, replay, self-healing, and what lives in .reticle/.

Handing results to something else

The CI artifact

Note the run wrapper. The artifact is nested one level down, which matters if you are parsing it in CI. A versioned, machine-readable record of a verification run. This is what you hand to CI, a dashboard, or a platform embedding Reticle in its pipeline. A passing flow additionally carries oracle, the consequence that made it a pass, so a green row says what it proved rather than only that it was green.

The capability contract

Read it back with reticle_capabilities { "fromDisk": true }, which needs no browser at all. Persists the app’s live capability registry to a git-checked file. A fresh agent can read it without booting the app, and a diff shows when someone removes a signal something depended on.

Talking to the human

narrate writes a line into the presenter panel, so the person watching knows what you are doing. review drains the bugs they pinned on elements from that panel:
An empty list here is a real reading, not a missing one.

Discovering the rest

Returns every tool with a one-line summary, and the profile currently in force. Load the full argument grammar for the ones you want before calling them:
Do this rather than guessing at arguments. Reticle refuses a call with unknown parameters rather than running it, and says why: “NOT applied, so any result would be an answer to a different question.” Good behaviour, but a wasted turn you can skip.
Last modified on September 23, 2026