Back to Whitepapers
The Sum of Your Gaps: What Fixed-Task AI Studies Do Not Measure
Whitepaper 10 min read

The Sum of Your Gaps: What Fixed-Task AI Studies Do Not Measure

Fixed-task studies measure performance within familiar work. A one-person repository history raises—but cannot answer—a different question about whether AI changes the range of work people attempt.

NC

Nino Chavez

Product Architect at commerce.com

Reading tip: This is a comprehensive whitepaper. Use your browser's find function (Cmd/Ctrl+F) to search for specific topics, or scroll through the executive summary for key findings.

Executive Summary

Two influential studies measure how generative AI affects performance within an existing kind of work. One studies customer-support conversations. The other studies experienced developers completing issues in repositories they already know.

These studies measure how AI affects performance within a fixed kind of work. They do not measure work someone only attempts because AI made it accessible. A repository history can demonstrate breadth in one person’s activity, but it does not prove that breadth predicts gain, quality, or value.

This paper examines that boundary through a census of 11,801 commits across 117 repositories belonging to one practitioner. The record contains activity across coding, design, operations, research, sales, documentation, and other categories. It supports a narrow observation: one person’s range of activity expanded beyond the work he was trained to do.

The record does not establish that AI caused the expansion. It does not evaluate the quality of the added work. It does not measure its economic or organizational value. It cannot support a formula connecting breadth to productivity.

Whether AI changes the range of work people attempt—and whether that added breadth creates value—remains an open research question.


1. The finding depends on the question

The following studies answer important questions about AI and productivity:

StudySettingOutcomeReported result
Brynjolfsson, Li, and Raymond, Generative AI at Work5,179 customer-support agents performing their existing jobIssues resolved per hour14% average increase; 34% for novice and lower-skilled workers; minimal effect for experienced and higher-skilled workers
METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity16 experienced developers completing 246 issues in repositories they knewTime required to complete an issueDevelopers took 19% longer when AI was allowed

The two studies differ in population, task, tool, and result. They share one feature: the participant remains inside a familiar kind of work.

The support agents continue handling support conversations. The developers continue implementing issues in repositories to which they have contributed for years. The treatment changes whether AI is available. It does not ask whether AI changes the set of work the participant considers attempting.

That is not a defect. It is the boundary of the research question.

The studies support claims about performance within their measured settings. They do not support a conclusion about all work a person might take on with access to AI.


2. Task selection is outside a fixed basket

A fixed-task study starts with a basket of work and compares performance with and without a tool. The basket must remain stable for the comparison to mean anything.

But a stable basket removes another possible effect from view: the tool might change what enters the basket.

Examples include:

  • a developer attempting a brand system instead of hiring it out
  • a founder writing internal documentation that previously remained unwritten
  • a designer building a working prototype instead of stopping at a mockup
  • an operator evaluating a dataset that previously required a specialist

The question is not whether these examples are common or valuable. No evidence presented here establishes either claim. The point is methodological: work newly attempted because a tool made it accessible has no observation in the no-tool condition.

METR encountered this problem in a later experiment. Its February 2026 update reports that 30% to 50% of surveyed developers withheld some tasks because they did not want to perform them without AI. Some developers also reported selecting different kinds of tasks or changing the amount of documentation and testing they produced.

METR concluded that these selection effects made its later productivity estimate difficult to interpret. That does not invalidate fixed-task research. It shows why task choice and task performance are different variables.


3. A one-person repository history shows breadth

The case examined here includes 104 GitHub repositories—47 public and 57 private—plus 13 local checkouts that had never been pushed. Together they contain 11,801 commits.

Each commit was classified using two signals:

  1. terms in the commit message
  2. paths of the files changed

A commit counts toward a category when either signal matches. Categories overlap by design. A payment-flow change can contain coding, sales, and operations work at the same time.

The resulting activity spans 16 seeded categories:

Work typeCommitsRepositories
coding7,839107
operations2,73089
tooling2,16184
research1,86681
run and maintain1,78387
strategy1,61264
design1,59088
sales1,40952
support83056
copywriting78568
photography67733
marketing67865
brand59370
distribution30241
decks20226
video19837

An open-vocabulary pass found additional activity the seeded list did not name, including methodology, prototyping, handoff, translation, evaluation, data work, architecture decisions, and audio. Another 845 commits remained unmatched.

The counts establish breadth in the recorded activity. The repository history also shows that this breadth emerged over time around a primary base in software development.

It is one case. It is self-selected, self-classified, and mostly private. It can demonstrate that the fixed-task frame excludes a real pattern in one person’s work. It cannot estimate how often that pattern occurs.


4. The census measures activity, not gain

The census has already produced one false result in the direction of its author’s argument.

The first analysis reported that Markdown outnumbered TypeScript. The collection script retained only the first 40 changed files in each commit. Large commits were code-heavy, so the cap removed more code than prose. The analysis also described repeated edits as distinct files.

Recomputed without the cap:

Distinct filesEditsEdits per file
Markdown7,40425,4243.4
TypeScript13,47834,0072.5
Files under docs/2,76617,6626.4
Files under src/4,12916,0003.9

TypeScript leads on distinct files and edits. The earlier claim is withdrawn.

The correction leaves two observations intact. The corpus contains substantial documentation activity, and prose files were revised more often than code files. Neither observation establishes gain.

A commit records that something changed. It does not establish:

  • that the work produced a finished deliverable
  • that another person approved it
  • that an independent reviewer judged it good
  • that it saved money or time
  • that it created value
  • that AI caused the work to happen

The distinction is fundamental. Activity is observable in the repository. Gain is a conclusion about a counterfactual: what would have happened without the tool. This census has no valid control condition.


5. Claim ledger

The evidence supports some claims and leaves others open.

ClaimStatusBasis
The cited studies measure performance within an existing kind of workSupportedTheir published methods and outcome measures
A fixed basket cannot directly observe a task absent from the no-tool conditionSupported as a property of the designComparison requires the task to exist in both conditions
One repository history contains activity across many kinds of workSupported for this caseAuthor-reported commit and path classification; most source records are private
The range of activity expanded in this caseSupported as a bounded case descriptionRepository chronology and the subject’s record; no population claim
AI caused the expansionNot establishedNo control condition or causal identification
The added work was goodNot establishedNo independent quality review
The added work created valueNot establishedNo outcome or economic measure
Greater breadth predicts greater AI gainOpen hypothesisOne case cannot establish a relationship

The final row is where the earlier version of this paper overreached. It represented gain as the sum of gaps between a person’s capability and an AI system’s capability across different kinds of work.

That may be a useful hypothesis. It is not a result. The census contains no measure of capability gaps and no measure of gain. The formula has therefore been removed.


6. What a useful study would need to measure

The open question is not merely whether AI makes people faster. It is whether AI changes the range of work they attempt, and what happens to the added work.

A study designed for that question would need at least five measurements.

6.1 The considered set

Record the work a participant considered before any assignment occurred. This includes tasks rejected as too expensive, too unfamiliar, or outside the participant’s role.

6.2 The attempted set

Track which tasks the participant actually begins with and without AI access. The difference measures task selection rather than task speed.

6.3 Completion

Separate an attempted task from a completed artifact. More starts are not automatically more gain.

6.4 Independent quality

Have qualified reviewers judge the outputs without knowing the tool condition. Quality must be assessed for each kind of work, not inferred from activity or self-report.

6.5 Value

Define the outcome before collecting data. Depending on the work, value might mean adoption, revenue, avoided cost, reduced time, fewer errors, or a stakeholder choosing to use the artifact.

The study would also need multiple participants with different levels of depth and breadth. That would make the central hypothesis testable:

After controlling for skill and context, does the number of newly attempted work categories predict independently judged gain?

Possible results include a positive relationship, no relationship, or a negative one if broader activity creates more low-quality output. All three remain consistent with the evidence presented here.


7. Conclusion

Fixed-task studies tell us whether AI helps someone perform familiar work faster. They do not tell us whether AI changes the range of work that person attempts.

This repository history shows that range expanded in one case. It does not show that AI caused the expansion. It does not show that the added work was valuable or good. It does not show that breadth predicts gain.

The title of this paper now names a question, not a formula. A person’s gaps may matter if a tool makes unfamiliar work accessible. We do not yet know whether the sum of those gaps predicts anything worth having.

That is the study still missing.


Provenance and limitations

  • Study results and methods were checked against the NBER paper page, METR’s early-2025 study, and METR’s February 2026 methods update.
  • The aggregate census figures come from an analysis of the author’s repository history. Most underlying commits are private, and the public claude-recall-cli repository does not currently contain the exact census scripts or dataset needed to reproduce these tables. Treat the figures as author-reported until those materials are published.
  • Keyword classification is crude and overlapping. It records material touched, not delivered outcomes.
  • Photography and video are undercounted because much of their output is not stored in Git.
  • The corpus contains one practitioner. It supports a case description, not a population estimate.
  • Most repositories are private. Readers can inspect the method and aggregate figures but cannot independently inspect most underlying commits.
Share: