The agent wrote “done”. The lawyer wrote forty-one comments.
Every new model moves the line, above all in code, where there is an answer to check against. In knowledge work it moves less: the most capable models still hand me work that is wrong, or blind to something I gave them, in the same even tone as the good work. This summer a chain of them delivered fifty-four articles as done; the fact-check step written into the chain never ran. A lawyer has read five pages so far and left forty-one comments, three of them the kind that cost a reader money. So the question is not whether agents will be wrong. It is what you would notice, when, and who.
The step that never ran
The first report I should not have believed came in the summer of 2025, when the coding agents were a couple of months old. For two weeks I worked on one system, and each day the agent told me it had committed the code, saved it to the shared history you go back to when something breaks. When the code vanished, the two weeks with no commits were the two weeks I had spent building it, and I had not looked either.¹ That particular problem is largely gone; today they commit before I ask, and sometimes before I want them to. The one underneath it is not. The report that the work was done came from the same machine that was supposed to have done it.
I am a partner in a real estate agency in Spain, the one whose strategy I wrote about in Part 3. The digital side is mine. This summer the job was a library of articles to bring in buyers from abroad: fifty-four articles in five weeks, thirty-three of them about tax and law. The Beckham regime, the flat-rate deal Spain offers people who move there for work; wealth tax; what a Dubai-based Briton owes where.²
The chain was written down before we started, and I expected it to carry the work most of the way. It ran on the newest model on the market, weeks old and much talked about, and the trust in what it could do on its own was high, mine included. Deep research on nine areas, three rounds of it, over a hundred and eighty raw source files behind it. An editorial-director agent and an SEO agent read the research independently and converged on a plan. One writing agent per article, nine in parallel, under an instruction I still think is right: write only from the material, cite every number, never present a forecast as a fact, flag every gap. And then, on paper, the step that mattered: a checker agent verifies every claim against the primary sources.
The checker never ran. The instruction was there, in the workflow the chain was built from. The coordinating agent, the one that ran the chain, dropped it, and its own log says why: minimal verification between batches, because its context, the working memory a model has for one session, was running short and it wanted the batches to keep moving. I only found that line when the lawyer's comments sent me back to the log. What ran instead was a consistency pass that caught one number wrong across the set, and a list of the gaps the writers had flagged themselves. No agent checked the basic description of the Beckham regime, because no agent had flagged it as a gap.
The library was built the way libraries like it are built: one base article per subject, then an angle for each kind of buyer, the same regime seen from the United States, from Scandinavia, from a Briton leaving the Gulf. It is a good structure for readers. It is also a copying machine. The Beckham sentence was wrong in the base article, and every angle inherited it. It was in eight articles.
No reader outside the company has seen any of it. The articles sit behind a PIN until our lawyer, an external partner who reviews everything the agency publishes on tax, has read them. He has read five of the thirty-nine pages I asked him to read, three of them articles, and left forty-one comments. Three would have cost a reader money. The Beckham article said the regime taxes only Spanish-source income, with most foreign income exempt. His correction: salary and director's fees are taxed whatever their origin, and gains on Spanish property or shares are taxed as for any resident. Someone moving on the strength of the article's sentence would have been planning around income the tax office counts. The wealth-tax article's opening says the region the agency works in has zeroed the tax; its own body, further down, describes the mechanism by which the tax still applies. And thirty-four pages are still in his queue.³
Here is the part I keep turning over. The correct description of the Beckham regime was in the research, twice, with citations, in the very material the writer was ordered to write only from: "employment income is taxed on a worldwide basis," one file says, and the other lists the exceptions line by line. The writer compressed the exception away and made the simplification the rule. In code, the answer to check against runs on its own. Here it sat in a folder, and nothing compared the article to it. And I did not read the articles, because I was not the right person to read them. I am not a tax lawyer, and I have never lived on that coast; my partners have lived there for years and sell property there, and the lawyer knows the law. My job was to build, with agents, a library of drafts that the people with that competence could review. I would probably have caught the Beckham sentence, because it contradicts how Spain taxes residents. I would not have caught all of it, I would not have known which ones I had missed, and the library would have reached the lawyer looking read. That is its own kind of done.
So the knowledge was not missing. The check was. And the chain reported done anyway, because "done" is what a chain says when it reaches its last step, whether or not the step before it ran.
It happens to the checkers too
The sharpest version I have comes from a legal knowledge system I build for a client. Every layer there has been tested since the first collection came back saying it had everything, and had not.⁴ The test I keep coming back to is this one.
Twelve agents, one jurisdiction each, identical task, identical report: done. Eleven of them were. The twelfth had invented seven of its sources, references that looked exactly like real ones and did not exist, and had read one provision backwards, writing that it limits a liability it in fact extends. Nothing in its report distinguished it from the other eleven. The only thing that did was a second agent, from a different model family, that took every citation back to the raw source. The rule the team wrote afterwards: the law itself is never fabricated, but the references around it sometimes are, and plausible references are the dangerous class.⁵
A lead-generation system I run tells the same story smaller. Its rule knowledge, what an installer of solar panels or heat pumps has to get right, had been built by an automated crawl and never checked against anything. Mapped this month, it described the same tax deduction in four incompatible ways. It had been live behind a client's chat agent, and any visitor who asked about the deduction would have got whichever of the four descriptions the retrieval happened to pick. That is the kind of thing you only find out by reading the logs. Rebuilding it took five passes, and the first checking agent approved two entries that repeated a rule a consumer ruling had overturned, because the ruling was not among its sources. You cannot validate against a source that was never fetched. No lawyer has read the new layer yet. The agents say it is done. I have learned what that means.⁶
What you would notice
I asked myself the question I would ask a client. A company of fifty puts an agent on customer service, or lets one quote prices. It gets something wrong. What do they notice first, and when, and who?
Honestly: for a long time, nothing. The error shows up when someone escalates. A customer who got the wrong answer and got angry. A price that did not match when the invoice came. By then it has already cost something, and the person who finds it is a person, not the agent. The agent was done weeks ago.
The research that measures this says the same, in two ways. A study this spring ran nineteen models through long editing workflows across fifty-two professions; even the frontier models, on average, had corrupted a quarter of the document by the twentieth exchange, and the damage was silent: sparse errors, severe, made without a word.⁷ Another looked at what agents write when they fail. In the coding runs it examined, three failures in four ended with the agent claiming success; in its customer-service runs, a second model set to read the agents' reports was fooled by the same confident closing.⁸
In our case the "who" was one lawyer with thirty-four pages in his queue, and the "when" was five weeks after the chain said done, when his first comment arrived; had the articles been public, the who would have been the first reader to act on the Beckham sentence.
So the owner's question is where the check sits, who owns it, and what happens where there isn't one.
How I place the check now
Customer service: I build a bank of the questions the agent is allowed to answer, from material a person has validated. If the answer is not in the bank, the agent does not answer; it hands the question to a human, and the human's answer goes into the bank. The bank grows, and over time the agent answers less from its own head and more from what someone in the company has signed off. It is no less capable, only less free. The number to watch is the share of answers that come from the bank. It should rise.
The number most companies watch is the other one. In early 2024 Klarna reported that its assistant had taken over two thirds of customer service chats in its first month, 2.3 million of them, and was doing the work of seven hundred people. Fifteen months later its chief executive said that cost had been "a too predominant evaluation factor" and that "what you end up having is lower quality."⁹ The quality had been checked, by the customers, later, at the customers' expense. Klarna could afford that. A company of fifty does not get a second year.
Prices: never an agent. The agent can talk to the customer and gather what they want. The number comes from a calculation the agent cannot touch, the same kind of thing a spreadsheet is, and it goes straight to the screen without passing through the agent. If the agent misunderstood the order, you can see it in the summary next to the number and fix it. If the agent had computed the number, you could not. The same holds for any figure a decision will rest on. The agent may describe it. It may not produce it.
Before any of that, sort. The old seven-question script that IT support used to run before they let you talk to anyone is the right idea. Automate the sorting and the known simple cases, the "here is a shipping label, send it back", and let the humans have the rest. On the cases you could not list in advance the agents are still wrong too often to leave alone, and the benchmarks that measure it say the same.¹⁰ Treat one as a junior who does most of the groundwork, with one difference. A junior's "done" you check because you know they are junior. The agent's sounds like a senior's.
And for anything that is written: a crew, not a reader. As soon as something exists I have it read from several angles, a fact-checker, an editor, the person it is meant for. Some of what comes back overlaps. Much of it is unique to the angle. I switch models between collecting and correcting, one gathers and another checks, sometimes a third from outside, because I have seen the check go both ways depending on who does it. I give the checker the lessons from last time, starting with the one from Part 2: search in the customer's language. Three more, from the rounds this summer: every fix goes back through review, because fixes introduce errors of their own; "not found" requires a documented search, because the file nobody searched is where the answer was; and you keep going until a round finds nothing new.⁵ And one more, from the library: the checker is the step a coordinating agent drops when its context runs short, whatever the instruction says, so I no longer check that it was ordered. I check that it ran.
The line that holds
Part 6 was about the question you give an agent. This part is about the report you get back. "Done" tells you when the agent stopped, not what state the world is in. The work can be delegated. "Done" cannot.
I use this technology every hour of every working day, and it has made building things cheaper and faster than I would have believed two years ago. What it has not earned is the walk away. Longer autonomy is the direction the companies that make the models are selling. On the evidence I have, it is the wrong variable. The right one is how long an agent runs before someone reads, and who that someone is. In my own ledger the difference is plain: fifty-four articles ran five weeks without a reading; the layer that held was read four times in a month, and the fourth reading found nothing.⁵
Three questions to take with you
- When an agent tells you it is done, how often do you go through what it actually did?
- Not its summary; the work. Open what it produced and compare it with the material it was told to work from. That is the check the chain in this article skipped, and the knowledge to catch the error was in the files the whole time.
- Which of your AI answers reaches a customer without passing through material a person has validated?
- Build a bank of the questions the agent may answer, from material a person has signed off, and route everything else to a human whose answer then goes into the bank. The number to watch is the share of answers that come from the bank. It should rise.
- Where in your business is an agent working out a number that a spreadsheet should?
- Take the number out of the agent's hands. The agent can gather what the customer wants; the figure comes from a calculation it cannot touch and goes to the screen next to the agent's summary, where a wrong reading is visible and a wrong calculation would not be.
Notes and sources
- The coding agent was Claude Code, released as a research preview on 24 February 2025 and made generally available on 22 May 2025. Public reports of false completion claims from that summer, both since closed, include issue #1501 in the tool's GitHub repository (2 June 2025: "Claude Code consistently reports having completed actions, tests, or operations that were not actually performed") and issue #4550 (27 July 2025), whose transcript has the agent say "I lied about adding those warning comments". In April 2025 the research group Transluce reported that a pre-release frontier model from another vendor "frequently fabricates actions it took to fulfill user requests", most often claiming to have run code it could not run.
- Fifty-four articles built over five weeks in June and July 2026 for the agency's advice section; thirty-three are to be published under the lawyer's review and will not go live until he has signed each one. The research: nine areas, three rounds, 182 raw source files. The pipeline as designed: plan, write, fact-check ("a checker agent verifies every claim against the primary raw sources"), SEO pass, editor sign-off. The pipeline log records that the fact-check step was not run: "minimal verification between batches", with the reason given as limited context in the main session.
- The lawyer's review, late August to early September 2026: forty-one comments on five pages of thirty-nine requested. The Beckham error: the article's "taxed on Spanish-source income only" against his correction that salary and director's fees are taxed "regardless of the origin" and Spanish-source gains "are taxed as a regular tax resident even under Beckham law"; the phrase appears in eight articles. The wealth-tax error: the introduction states a regional exemption while the body describes the reinstated tax. A third comment concerned which tax treaty, if any, applies to a Briton resident in Dubai. The correct Beckham description exists in the project's own research files from June 2026: "employment income is taxed on a worldwide basis", and a second file listing the exceptions with sources.
- Part 2, "When in Rome, search in Italian", for how the gaps in that first collection, a swarm of agents sent out for every local law under one international tax framework, were found, beginning with an Icelandic ruling the agent could not find until it was asked whether it had searched in Icelandic.
- From the project's own verification records: a twelve-jurisdiction test in August 2026 (seven fabricated secondary references and one inverted provision in one jurisdiction; validation by a second model family, subsequently made mandatory) and a four-round review of a 168-item corpus the same month (24, 3, 1, 0 findings per round; two round-two errors introduced by round-one fixes; one round-four correction of a round-three reviewer's own error; round one had approved items marked "no position on record" where the position existed in an unsearched file). The three rules in the text come from that review. Client and system not identified.
- The system's rule-knowledge layer, rebuilt in September 2026: 36 cells over nine rule areas and four product lines, 105 raw sources, collection by one model, validation by another, a gap-closing pass that closed three cross-cutting gaps with eighteen of those sources, an update of 16 cells, and a second validation of 160 claims with none unsupported. Three previously assumed rules corrected against statute or a ruling; two cells repeating one of them had been approved by the first validator. The legacy layer it replaced: 177 claims mapped against the new structure, 15 of 36 cells without coverage, two statutory references in the whole set, four mutually incompatible descriptions of one deduction, and authority URLs that look constructed. No human lawyer has reviewed the new layer; validation is by agents against primary sources. Customer not identified.
- Philippe Laban, Tobias Schnabel and Jennifer Neville, "LLMs Corrupt Your Documents When You Delegate", arXiv 2604.15597, 17 April 2026. DELEGATE-52: long delegated editing workflows across 52 professional domains, 19 models. "Even frontier models (Gemini 3.1 Pro, Claude 4.6 Opus, GPT 5.4) corrupt an average of 25% of document content by the end of long workflows"; "agentic tool use does not improve performance"; the errors are "sparse but severe" and "silently corrupt documents, compounding over long interaction". In the paper's Table 1, one model, Claude 4.6 Sonnet, scores 92.2 after two interactions and 66.0 after twenty; the three frontier models average 75.2 after twenty.
- "From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents", arXiv 2606.09863: "LLM agents can fail silently by asserting task completion when the environment state shows otherwise"; among self-assessing coding-agent trajectories with explicit status claims, 75.8 percent of failures were false successes; "LLM judges fail reliably: no configuration across 5 judges, 5 prompt strategies, and full task specifications exceeds AUROC 0.65 on tau2-bench", the customer-service benchmark, because the judges "rely on surface completion proxies", above all "confident closing language"; on the coding benchmark, where agents write no closing message, the judges reach only 0.54.
- Klarna, February 2024, as reported by Customer Experience Dive: the assistant "had taken over two-thirds of customer service chats" in its first month, "2.3 million in total", average resolution under two minutes; the claim that it did the work of 700 representatives is reported in the same article. Sebastian Siemiatkowski to Bloomberg, 8 May 2025, as reported by Customer Experience Dive, 9 May 2025.
- Two benchmarks. TheAgentCompany (Carnegie Mellon University, arXiv 2412.14161): 175 workplace tasks in a simulated company; in the December 2024 version "with the most competitive agent, 24% of the tasks can be completed autonomously", in the September 2025 version "the most competitive agent can complete 30% of tasks autonomously" (30.3 percent, 39.3 percent with partial credit). CRMArena-Pro (Salesforce, arXiv 2505.18878, May 2025): "leading LLM agents achieve only around 58% single-turn success on CRMArena-Pro, with performance dropping significantly to approximately 35% in multi-turn settings". Both measure whether tasks were completed, not whether the agents reported them as completed.