Back to all essays
AI & Automation • • 7 min read

Rigorous About the Wrong Question

I asked an AI agent to audit my app for the person it was built for. It sent five sub-agents to check documents and none of them looked at the screen.

NC

Nino Chavez

Product Architect at commerce.com

I asked an AI agent to audit my app. The agent split the job across five AI sub-agents. Each sub-agent got a written instruction sheet, which I will call its brief. Eleven screenshots of the app, taken on a real iPhone, were sitting in the code repository the whole time. None of the five briefs said “look at them and judge what you see.” The report that came back was careful and well sourced, and it answered a question I had not asked.

Two rules came out of this.

First, when the question is about what a person sees on a screen, look at the screen yourself before you send anyone else to look. A reviewer only sees what its brief asks about.

Second, before the briefs go out, read them against your original question. If you asked mostly about the user and the briefs are mostly about documents, you can see the mismatch in thirty seconds. After the briefs go out, it takes a correction to find it.

What I asked, and what the reviewers were told

Minder is a day planner I have been building for one person. Zoey manages a household calendar she does not own and cannot edit. Her day is the thing the app is supposed to make easier.

I asked the agent four questions. How far has the app drifted from its mission? Is every interaction right for what Zoey needs to do, see, and know? What does she do often, and is that quick? What is repeatable, automatable, saved and reusable?

One of those is about process. Three are about her. The word “right” asks for a judgment. It does not ask for an inventory.

The agent wrote five briefs. One sub-agent mapped code folders to the decision records that authorized them. A decision record is a short document that says what was decided and why. One sub-agent compared the design rules against the code. One listed which features could be reused or automated. One read the screenshots, but only to check them against the design rules. The fifth was named “user fit,” and its brief asked for tap counts. Not one brief asked whether the screen was right for Zoey.

The report was accurate

Every finding held up on a second check. The only recorded verdict from Zoey was a rejection of the first build, on 2026-08-23. The newest build had zero installs. One decision record said not to turn on a feature in production and declined to set a timing value for it. A code change the next day set the value and turned the feature on. Five feature switches were on in the release build while three documents said they were all off.

Those are real problems. Several of them block a release. They are also all about process.

I read the report and sent back one sentence: “this report feels like it just audited the mechanical parts and none of the application functional parts for a user.”

What the screenshot showed

That sentence sent the agent to the screenshots. It opened them itself and wrote down what it saw.

One screenshot shows a busy day. That is the day the app exists for. The screen has a card for what is happening now. Below that card sit two buttons for adding things, a strip of dates, a section heading, and a filter control. Only then comes the card for what is happening next. That card starts at the bottom edge of the screen. Its start time is cut off mid-line.

The app’s own design rule says the screen must answer “what is now, what is next” within five seconds. On the day that matters most, it does not answer the second half at all.

Another screenshot shows an upcoming gymnastics class. Under the class name and its start time, the card asks for a map location, spends three lines explaining why, and offers a large button to go pick one. The child’s name, Maya, comes after all of that, in small grey text.

None of that needed a specialist. It needed someone holding my question to look at the picture.

Why a reviewer can look and still not see

The sub-agent that read the screenshots did its job. Its brief said to check them against the design rules, and it did. Its brief did not say “is this what Zoey needs,” so it had no reason to report that a child’s name had lost to a map prompt.

Delegating a search works. An agent sweeps a hundred files and the answer comes back intact. Delegating the looking does not work the same way. Looking is only worth something when the one doing it holds the question.

The briefs came out process-shaped because the repository is process-shaped. Minder has seventeen decision records, a table tracing every requirement to its evidence, and a written definition of done. All of that is easy to search, and searching rewards process questions. The user side was there too. The requirements document names forty-three things Zoey needs to do, numbered FP-0 through FP-42. “Open the destination” and “Set a leave reminder” are two of them, and they are the gymnastics card. The agent found that list after my sentence, not before.

The app’s own review had the same blind spot

Two days earlier, a separate review of the app’s screens had reported that four screenshots showed a “Mark done” button on read-only calendar events. A decision record dated 2026-09-01 had already removed that button from calendar events, and the code change landed the same day, before those screenshots were taken. The screenshot of a calendar event shows a map button, a details button, and no “Mark done.”

The review was quoting a design document that still described the old defect. The document and the review agreed with each other. Neither matched the picture. A reviewer can have a screenshot open and still be reading a stale document.

What I cannot claim

The agent’s closing report offered me a line: the team wrote 96,000 lines of code and never put the product in front of the one person it was built for. The line count is right. The rest is not. A test build went onto Zoey’s phone on 2026-08-25, and she has had a later build since. What is true is narrower. Her only recorded verdict is that first rejection, and every check added since then measures what the team could verify without her.

I also cannot say which fix worked. After my correction, the agent changed two things at once. It rewrote the briefs around Zoey’s day instead of around the code, and it told every sub-agent to look at the screenshots before reading anything else. The second round found more problems of the same kind. I do not know which change did it.

The two rules, in the words I now use

When the question is about what someone sees, the person or agent forming the verdict opens the screen first. Reading a summary of what a screen shows is not looking at it.

Before sending briefs to reviewers, sort your question into kinds. Some parts are about process. Some are about what the user sees and does. Then read the briefs and count which kind each one serves. If the counts do not match your question, fix the briefs. That is the only cheap moment.

The audit was rigorous. It was rigorous about the wrong question, and rigor did not tell anyone.

Share:

More in AI & Automation