The four failures a screenshot approves
Mock data
POST returns
200, the row appears, nothing persisted. Reload and it’s gone. The screenshot
during the demo is perfect.Dead handler
Button has an
onClick, the onClick calls a function, the function was never wired to the
store. Looks connected all the way down.Double submit
One click, two POSTs, customer charged twice. Zero pixels differ.
Silent validation
The form accepts
"abc" as a quantity and sends it to your database. Renders beautifully.”Just use a better model”
This is the intuition worth killing, because it sounds reasonable and costs teams months. Model quality is not the constraint. You cannot see a network call. You cannot photograph a state mutation. A screenshot of a page whose bug is non-visual contains no signal about that bug, so asking a stronger model to look harder produces a more confident version of the same wrong answer. Improving the model improves the reading of the evidence. It does not create evidence that was never captured.What “reading the program” means instead
Reticle runs inside the app and reads what actually happened:net.total: 1 kills double-submit. stateDiffs kills mock data, the store genuinely changed. signals is the app declaring success in its own words. consoleErrors: 0 kills “it worked but threw”.
No pixel in the world carries that.
A real one, caught on the first run
On the first drive of a real production dashboard, before any instrumentation, Reticle’s network observation flagged two endpoints returning500. A database migration that had not been applied.
The page rendered fine. The sidebar loaded. Nothing looked broken. A screenshot agent would have called it done and moved on.
That is the thesis in one incident: “looks done” and “is done” are different properties, and the difference is usually non-visual.
When a screenshot is the right tool
Screenshots are not useless. They are the correct instrument for a specific class of bug, and this page would be dishonest without saying so:- A font that failed to load
- A GPU or compositing glitch
- Layout that breaks only at a particular viewport
- Anything where the actual rendered frame is the thing under test
reticle_screenshot and reticle_visual_diff for exactly this, and if pixel regression is your primary concern, a dedicated visual tool will serve you better.
The argument is not “screenshots are bad”. It is that a screenshot cannot answer “did the code work”, and that is the question an AI coding agent needs answered on every change.
The honest limits
We have not proven that Reticle makes agents fix more bugs. A controlled test did not show a fix-rate improvement; the benchmark was confounded and we are still working on measuring it properly. What is measured is detection: 10 of 10 injected regressions caught with zero false alarms, against 9 and 8 for the alternatives.FAQ
Would a better vision model fix this?
Would a better vision model fix this?
No, and this is the single most expensive misconception on the page. Model quality improves the
reading of evidence; it cannot create evidence that was never captured. A screenshot of a page
whose bug is non-visual contains no signal about that bug at any resolution, so a stronger model
produces a more confident version of the same wrong answer.
So screenshots are useless?
So screenshots are useless?
No. They are the correct instrument for a specific class of bug: a font that failed to load, a GPU
or compositing glitch, layout that breaks only at one viewport, and anything where the rendered
frame is the thing under test. Reticle ships
reticle_screenshot and reticle_visual_diff for
exactly those. If pixel regression is your primary concern, a dedicated visual tool will serve you
better.What does Reticle read instead of pixels?
What does Reticle read instead of pixels?
The program: the network log (did the request fire, how many times, what did it return), the
console, client-side routing, registered store state, and any signal the app itself emits. A
response carrying
net.total: 1 rules out a double submit; a non-empty stateDiffs rules out
mock data; consoleErrors: 0 rules out “it worked but threw”.Does that mean Reticle catches every bug?
Does that mean Reticle catches every bug?
No. It has not been shown to make agents fix more bugs; a controlled fix-rate test did not show
an improvement and we consider that question open. What is measured is detection: 10 of 10
injected regressions caught with zero false alarms, against 9 and 8 for the alternatives. And a
purely visual regression still needs the opt-in visual diff.
Can I keep using screenshots alongside Reticle?
Can I keep using screenshots alongside Reticle?
Yes, and for a visual product you probably should. The argument here is not that screenshots are
bad, it is that a screenshot cannot answer “did the code work”, which is the question an AI coding
agent needs answered on every change.
See a real verdict
Five minutes, and the output on that page was captured live.
The full benchmark
Method, raw numbers, and where Reticle comes second.