← All posts

AI tester over 11 weeks: 77 bugs from the agent, 16 from the human

A Claude Code session diary: 277 agent hours, 94 confirmed finds, where a reference pays off, and why an agent hour does not replace a human hour.

AI tester over 11 weeks: 77 bugs from the agent, 16 from the human
Contents

In brief

On Habr, a tester published an eleven-week diary of working with Claude Code as an AI tester. Session transcripts show 277 agent hours and 213 hours of human presence, with 94 confirmed findings — 77 from the agent, 16 from the author. The picture is closer to doubling throughput than to a “ten times faster” promise: every agent hour still cost about three quarters of a human hour for setup and review.

What happened

The author maintains an open-source rule pack, paranoid-qa, that pushes the agent to test with evidence instead of tidy reports. This piece is not the method — it is the ledger: where the hours came from, who found each defect, and what happened to each candidate. Agent hours are summed from active intervals in session logs; pauses longer than thirty minutes are dropped. Wall-clock from open to close is useless: a session left open for a week will report hundreds of hours for an hour and a half of work.

Defects are stricter. Only findings confirmed by the final session verdict count. Mid-session agent lists nearly doubled the picture: about eighty finds mid-session shrank to about forty on final verdicts. Candidates die under the agent’s own later checks and under human review. Roughly one hundred thirty candidates never reached confirmation; about forty-five left as questions for another team’s ownership.

The hunting grounds barely overlap. The agent wins where you need systematic state coverage or reading what a human will not read: a dropdown value leaving as [object Object] in the request body while the screen looked fine, a 500 instead of a validation error, a startup race on a widget, a code-slice mismatch between test and production that never shows in the browser. The human catches what only shows up if you are present in the live product: horizontal scroll on a phone, a bare error page when the server blinked, layout drift on a double open. The useful joint mode is human instinct plus the agent’s shovel: one noticed the page jumping, the other traced the carousel height recalculation in twenty minutes.

Why it matters

The richest hour is the one with a reference: a mockup, a spec, a production contour, or a failing automated test. Comparing implementation to a reference yielded more than one find per hour; ordinary UI passes sat around 0.6. It is like checking a delivery against a packing list: with a list, sorting the crate is fast; without one, you wander the warehouse and hope your eyes catch the bruise. Test design shows zero product defects in the table — a misleading row, because cases written there later feed the runs upstairs.

Discipline kills some false positives, not all. The agent reported thirty-one broken links, then walked them with a proper Accept header and cleared all thirty-one before filing. In another case it stuffed an internal-error body into its own network stub, saw that text on screen, and filed it as a find — testing the echo of its own mock. That day the pack gained a control question: did the system produce this value, or did I? Six misses against seventy-seven agent finds were caught by the human; the last two were on retest, when the agent stared only at the fixed spot and did not re-walk the whole block.

The cost story also breaks the slogan. Roughly a hundred dollars a month for about one hundred ten agent hours puts an agent hour near eighty rubles. Comparing that to a staff engineer’s hour “fifteen times cheaper” is wrong: every agent hour still took 0.77 human hours. On the same tasks the saving lands around sixteen percent. The larger win is work that would never have been scheduled: line-by-line slice diffs, forms with third-party scripts down, dozens of links across browsers.

In practice

This is a field diary, not a drop-in process for every team. Steal the counting method and the rules; do not copy the 77-to-16 ratio as destiny.

  1. Measure active time from session transcripts; open-to-close wall-clock lies by tens of times.
  2. Score only final verdicts; mid-session defect lists are nearly double reality.
  3. Give the agent a reference — mockup, spec, production, failing test — and ask it to compare, not “poke the form.”
  4. On network stubs, check UI behavior, not the texts and codes you planted in the response.
  5. Personally re-check Fail results from sub-agents; after fixes, walk the whole block visually, not only the patched spot.
  6. Price the work as your time share plus agent hours, not as “the agent replaced an engineer.”

Takeaway

Over eleven weeks the AI tester did not let the human “set it and walk away.” It doubled the volume of work that has a reference and a long, monotonous search space, and left presence in the live product, deduplication, and the final verdict to the person. The multiple-speed promise does not show up in these numbers; what does show up is that without counting discipline and verification rules, an agent will happily report a harvest that is not there.

Comments

Loading comments…