I Planted Eight Bugs and Measured What an Autonomous QA Agent Found

Two runs of the same product, against the same application, at the same URL, returned different defects. Precision was flawless both times. Recall behaved like a sample, not like a measurement.


Disclosure, before anything else

I sell independent AI red teaming. Replay QA, the product measured below, is a competitor of mine. You should read everything that follows with that in mind.

So here is what I did to keep myself honest. I wrote the test application before I ever ran their product. I wrote down every defect and every trap in advance. I am publishing the numbers that make them look good alongside the ones that do not. At the end I list the mistakes I made while doing it. And the application is published in full, so anyone, including Replay, can repeat this measurement and contradict me.

Why this is not another "I found a bug in X" post

Almost everything written about AI testing tools is anecdotal. Someone runs a tool against a real application, the tool finds something, and the write-up says the tool is good. Or it misses something, and the write-up says the tool is bad. Neither tells you the thing you actually want to know, which is: out of everything that was wrong, how much did it find, and of what it reported, how much was real.

You cannot compute that against a real application, because nobody knows the full list of what is wrong with a real application. So I wrote the application myself.

The test application

A single-page shop called Harbour Supply: catalogue, product page, cart, checkout, confirmation. No randomness, no dates, no storage, so every session starts identically and runs are comparable. Zero console errors.

I planted eight defects:

  1. Cart arithmetic. Above a quantity of two, the line is priced as unit price times quantity minus one. Three notebooks at $10.00 give a line total of $20.00.
  2. Price mismatch. A lamp is $19.99 in the catalogue and $24.99 on its own page. The cart takes the higher one.
  3. Email validation does nothing. The string abc is accepted and the order goes through.
  4. Dead control. The promo code Apply button has no handler. Nothing happens, no message.
  5. Broken link. The footer Shipping information link points at a page that does not exist.
  6. Stale state. Removing the last item empties the cart table but leaves the header badge unchanged.
  7. Required field is not required. The delivery address is marked required and the form submits without it.
  8. Missing accessible name. The card number input has no label, while the CVC field next to it has one.

And five traps, which is the part most write-ups skip. A trap is correct behaviour that looks like a defect: a submit button disabled until two fields are filled, an unknown promo code rejected with a clear message, an external link opening with rel="noopener noreferrer", a cart spinner that shows for about 0.7 seconds, and quantities above ten clamped with an explanation.

Traps are how you measure false positives. Without them, precision is not a number, it is a hope.

Making the ground truth trustworthy was harder than building it

Ground truth only works if the application contains nothing wrong except what I planted. My first build failed that test twice.

A quantity input stretched to 328 pixels instead of three characters wide, because an input[type=text] selector outranked my class. And the cart table overflowed the viewport at 390 pixels with no scroll container. Both are real layout defects. An agent reporting either of them would have been right, and my precision score would have counted it as a false positive, because it was not on my list. I found them with a separate pass across three widths and fixed both before the first run.

The second problem nearly destroyed defect five. Locally, the missing shipping page returned 404. On the live host, with no 404.html in the deploy, unknown paths were served the catalogue instead, the way a single-page app would. The planted broken link existed only on my machine.

The hosting is part of the application. Ground truth has to be verified at the address the agent will actually visit, not the one you developed against.

Run one

Eighteen journeys, six bugs filed.

Recallfour of eight
Precisionsix of six
Traps reported as defectszero of five
Duplicatesone

It found the broken email validation, the broken footer link, the stale cart badge and the missing input label. It missed the cart arithmetic, the price mismatch, the dead Apply button and the empty required address. Every single statement it filed was factually true about the application. Not one piece of correct behaviour was reported as a defect, and it explained in its own write-up why the disabled submit button was correct, without being asked.

The interesting part is not the score, it is the miss. Both defects involving money were missed, and one of them was missed inside a journey the agent had written for itself. Its own journey description said, quoting their output: "A cart that mis-sums totals, fails to update the badge, or cannot empty itself is surfaced as a defect." The cart does mis-sum totals. It found the badge problem in the same journey and not the arithmetic.

Why: the step raised the quantity to two, and my defect starts at three. The expected result was phrased as "the cart total has grown accordingly (roughly doubled for two units)". A check that stops one unit short, written as an approximation rather than an equality, walks past an off-by-one undercharge.

The price mismatch was missed differently and more interestingly. The agent opened the exact product with the mismatch, added it to the cart, and checked that the cart total matched the price it had seen. It did. The cart is internally consistent with itself. Nothing in the journey ever compared two screens, so a number that is wrong on one of them is invisible.

Run two

Same URL, same settings, same application, two days later. One exploration, twenty journeys, five new bugs.

It found the cart arithmetic, the price mismatch and the dead Apply button. Exactly the three that run one had missed.

Cumulative recall across both runsseven of eight
Precision, run twofive of five
Traps reported as defectszero of five
Never found by either runthe empty required address

It also filed two defects I had not planted, and both are real. The confirmation page claims a confirmation email has been sent, and nothing sends email. The confirmation route is unguarded, so navigating straight to it shows an "Order placed" screen for an order that never happened. Those are true findings about my application that I had missed myself.

One caveat I owe them and myself. It found the cart arithmetic but described it as "quantity input change does not recalculate line total/subtotal/total". The application does recalculate, just wrongly, by one unit, above a quantity of two. The observable symptom is exactly what it saw, because going from two to three does not change the displayed number. It caught the anomaly and misdiagnosed the mechanism. That is not a clean find, and it is not a miss either.

The result that matters

Two runs of the same product, against the same application, at the same URL, with the same settings, returned different sets of defects. Neither run alone was complete.

Precision was flawless in both runs. When this product tells you something is wrong, it is right. Eleven findings across two runs, every one of them true, zero traps taken, nothing invented. That is genuinely hard and most tools cannot do it.

But recall behaved like a sample, not like a measurement. If you run an autonomous tester once and it comes back with four defects and a set of passed journeys, you do not know whether you are looking at four defects or eight. In run one, a buyer would have received a green result on a cart that quietly undercharges, inside a journey that promised to check exactly that.

For a product sold as autonomous testing, this is the number that should be on the box, and it is the number nobody publishes: not how much it finds, but how much the answer moves between runs.

What held up

Said plainly, because leaving it out would be dishonest. Precision across eleven findings with zero false positives. Zero of five traps taken, twice. Two real defects found that I did not plant and had not noticed. Root cause analysis that was actual analysis: for the validation defect it read the page source, found the form marked novalidate, quoted the single length check in the handler, and correctly explained that novalidate disables the browser's own format checking. Evidence with timestamps, DOM values captured at the moment of the click, screenshots. And domain ownership verification that is actually enforced before every run, not just claimed.

What I got wrong

My first build of the stand contained two unplanted defects, which would have turned two correct findings into false positives and corrupted the precision number. I caught them, but only because I ran a separate check, and I nearly did not.

I built and verified the ground truth locally and almost measured against an application that differed from the one the agent would see. One of my eight planted defects did not exist at the live address at all.

I also ran only two runs, because the free tier gives 25 credits a month and the two runs cost 8.5 and 10. A third run would tell you whether the union keeps growing or settles, and I cannot tell you that. If the union keeps growing at run three, the honest reading of this whole piece changes, and I would rather say that than pretend two points describe a curve.

What this does not show

One application is one sample, and I built it around the defect classes I care about, which is not a neutral choice. The application is static, with no backend and no API, so the security portion of the product had nothing to probe and its empty result says nothing about that capability. Two runs cannot tell you the shape of the distribution, only that it is not a point. None of this is a verdict on the product. It is a measurement of one product against one application on two days.

Reproducing this

The full source of Harbour Supply is published here: harbour-supply-stand.zip. It contains the eight planted defects, the five traps, and the 404.html that the fifth defect depends on.

Deploy it anywhere, point an autonomous tester at it, and count. If your numbers differ from mine, I would like to know. Do not deploy it under a domain you use for anything else: an autonomous tester will follow links out of it.

Written by Nikita Kalachev, QA & LLM Evaluation Engineer at Platilus LLC. I do independent AI red teaming.
← Back to Research

This is what I do for other people's products

Independent AI red teaming. I test a production AI feature by hand, write down what it did, and deliver a dated report plus a one-page attestation you can attach to a security questionnaire. Three products, fixed prices from 700 USD.

Send the URL, get a scoped reply →