The Audit Reproduced the Failure
A small inspection of one runaway session became repeated parallel work before it had earned that expansion. The 119-of-118 report was the alarm.
Nino Chavez
Product Architect at commerce.com
The inspection burned through session allowance before it answered the question that started it.
It began as a narrow request: inspect one runaway session and find the work that had backtracked, regressed, or needed to be re-steered. The run expanded into parallel reading and verification work. Parts of that work were then stopped and restarted. By the time the audit produced a report, it had consumed the intraday allowance, about a fifth of the weekly allowance, and 40% of a model-specific allowance.
The result looked tidy: 119 of 118 findings verified.
That impossible number was not a typo to clean up. It was the signal that the inspection had grown expensive before it had earned the right to grow. A parallel audit needs a small first pass, a budget and concurrency cap, and a human checkpoint before it fans out again.
The first report could not count its own work
The audit had a journal. Each agent attempt could append a row. When a run was stopped and restarted, the new attempt wrote another row for the same candidate event.
The status report counted those rows as completed work. That is how it reached 119 verified findings inside a set of 118. Verification was still running.
An attempt, a candidate event, and a verified finding are not the same thing. The report had promoted one into the next without a stable identifier tying them together.
| Artifact | What it can establish | What it cannot establish |
|---|---|---|
| Agent transcript | An agent was started and produced text | The task’s conclusion is correct |
| Journal row | An attempt recorded an event | The event was counted once or finished |
| Tool exit code | A command returned successfully | The intended runtime property is true |
| Verification record | A claim was checked against named source material | The entire workflow is complete |
The system had records of all four types. A row showing that an attempt ran became evidence that a finding had been verified.
Parallel work earns its next round by changing the answer, not by producing more activity.
The counting repair is mechanical: count stable event IDs, not journal rows. Keep retries because they explain cost and failure history. Do not let them become extra findings.
Repeated fan-out hid behind a clean receipt
The final receipt described the logical run that finished. It did not describe the investigation that produced it.
Before that receipt existed, work had been launched, cancelled, and launched again under different settings. Some verification work ran again after those changes. The abandoned attempts still used time and allowance. They were simply absent from the summary.
That is the dangerous shape. A small inspection becomes an expensive parallel operation through a series of locally reasonable decisions: add more readers, rerun the uncertain slice, restart the cancelled attempt, then produce a final receipt that only represents what survived to the end.
The receipt made the work appear bounded because it only knew about the branch that finished. The actual resource constraint had already been hit.
This is not an argument against parallel work. Parallel work can be the right tool when a larger sample or an independent check will change the decision. It is an argument against treating expansion as free until a dashboard says otherwise.
The first pass should have had a fixed budget and a small concurrency cap. The next pass should have required a checkpoint: what did the first pass establish, what remains uncertain, and why will more parallel work change the answer? If a human cannot state that reason, the work has not earned another round.
The 119-of-118 alarm led to a second problem
The impossible total was enough to stop treating the report as a verdict. Inspecting the underlying records exposed a separate failure: the verification sample was biased.
Reader agents had labeled candidate events in three groups: the written guidance would have prevented the event, might have helped, or would not have helped. The labels were useful for deciding where to look. They were not evidence yet.
The verification rule checked every candidate in the “would prevent” group. It checked only portions of the other two groups.
| Reader label | Candidates | Transcript-verified |
|---|---|---|
| Would prevent | 44 | 44 |
| Would partially prevent | 248 | 70 |
| Would not prevent | 217 | 55 |
The fully checked group shrank from 44 candidates to 35. The lightly checked groups retained most of their original labels. Comparing those group sizes afterward would make the selection rule look like a finding.
It was not. The audit had not tested comparable samples under one decision rule.
The 35 events do survive direct transcript review. They are examples of failures the written guidance would clearly have prevented. They are not a population rate, a claim about every session, or a comparison across the three labels.
That boundary matters because a compelling rate is easy to repeat after the record that created it has disappeared. The alarm was not merely bad arithmetic. It was a report claiming more certainty than its own method could support.
Thirty-five examples survived. The rate did not.
The failed comparison does not make the extraction useless. The readers pulled a scattered session history into a record that can be inspected. They surfaced real events.
Thirty-five preventable events survived direct transcript review. That is a usable floor for the narrow claim: these examples happened, and the written guidance clearly addressed them.
Nothing in this record supports a population rate. Nothing here compares one model, prompt, or orchestration system with another. The process that extracted candidates also designed the filter, interpreted the results, and reported the rate. That is too much authority in one loop.
The work separates cleanly when each stage has a different receipt:
extract → candidate event, source location, and producer
count → stable event ID and deduplicated attempt history
verify → named source, decision rule, and surviving claim
Those roles can be performed by people, programs, or different systems. The separation is what matters. A process that designs its own filter does not get to certify the rate that filter produces.
Expansion needs a stop rule
The next inspection does not begin by filling the available parallel slots. It begins with a bounded first pass.
Set the budget before the first task starts. Set the maximum number of concurrent workers. Define the question the first pass must answer. Then stop.
Expansion requires a checkpoint with a human reason. That reason should be concrete: the first pass found a contradiction that needs independent review; a missing slice could change the decision; a new sample will test the same rule across groups. “More coverage would be nice” is not a reason to consume another round of allowance.
The receipt comes after that control, not before it. For every conclusion, it should preserve:
- the candidate event and its source location
- the stable event ID and deduplicated attempt history
- the source checked, the decision rule applied, and the claim that survived
If one of those pieces is missing, the conclusion stays a candidate. It may still guide the next question. It cannot become a rate, a completion claim, or a reason to spend more.
The first inspection became expensive before it became trustworthy. The next one needs to be able to stop.