← All articles Part of: AI Governance & Oversight

The Verification Step Nobody Owns

9 min read
Hand-drawn ink-and-coloured-pencil editorial illustration on cream paper: a thick stack of report pages seen from above and slightly to the side, its top page ruled with steel-blue bar charts, a pie chart and columns of text, with amber tabs marking pages along the stack's edges; a large black-handled magnifying glass lies across the front of the stack, and inside the lens the neat page edges resolve into a jagged crack and a spill of crumbling amber and cream fragments, with loose scraps scattered on the desk beyond it and a dark blue pen with gold trim resting to the right.

Key takeaways

  • Deloitte, KPMG and EY each shipped a report with fabricated citations inside roughly eight months. Three incidents in three firms that sell credibility is a pattern, not bad luck.
  • Only 5 of KPMG's 45 citations pointed correctly to the source they named, and 19 of EY Canada's 27 references were hallucinated, a 70% rate inside a 44 page report.
  • The failure was economic rather than moral. Verifying dozens of citations costs about what the original research costs, so the check was never funded and never happened.
  • One fabricated citation was inherited from an earlier blog post that invented the same source, and AI tools later quoted the flawed report back to users. Unverified claims circulate.
  • Price the verification before approving the work and name the person who owns it. An unowned check is not a control, and a review process can pass while verifying nothing.

In roughly eight months, three of the largest professional services firms in the world published reports containing citations that did not survive checking.

Not one of them was a small firm without process. All three sell judgment for a living, and all three have review layers designed to catch exactly this.

Start with the facts, because the interpretations that followed were mostly wrong.

In October 2025, Deloitte Australia delivered a report on welfare compliance to the federal Department of Employment and Workplace Relations under a contract executed in December 2024 worth about A$440,000, roughly US$290,000. The report contained references to academic papers that did not exist and a fabricated quote from a federal court judgment. Chris Rudge, a researcher in health and welfare law at the University of Sydney, flagged it. Deloitte refunded A$97,000, about US$63,000, which is less than a quarter of what the department had paid. The revised version added a disclosure that a generative AI system, Azure OpenAI, had been used in preparing it.

In June 2026, KPMG withdrew a global report titled “Total Experience: Redefining Excellence in the Age of Agentic AI,” which it had published in October 2025. The AI detection firm GPTZero examined its 45 citations and found that only five correctly pointed to the source they named. The rest, in GPTZero’s description, ranged from mangled and misleading to partially fabricated or too vague to verify. UBS, the UK’s National Health Service, Swiss Federal Railways and Transport for London all disputed claims the report made about their own AI deployments. The report also contradicted KPMG’s own concurrent CEO Outlook, citing 55% of CEOs prioritising AI investment against the 71% in the firm’s parallel survey. KPMG said it was reviewing the circumstances surrounding publication and that it expects its people to follow its guidelines on responsible AI use, “including human oversight.”

And in a 44 page report called “Points of Attack: Uncovering Cyber Threats and Fraud in Loyalty Systems,” credited to two partners and a senior manager, EY Canada cited 27 sources in its resources table. GPTZero’s investigation, published 14 May 2026, found that 19 of them were hallucinated, a rate of 70%. The fabricated references were attributed to real and respectable publications: BleepingComputer, Wired, Gartner, Forbes, McKinsey, Cisco Talos, TechCrunch. Inside the same document, the value of the loyalty market appeared twice with two different figures. EY Canada withdrew the report and said it was reviewing the circumstances that led to publication.

Three firms, three withdrawals, one refund. That is where most of the commentary stopped, and stopping there is what makes it useless.

The tally is the least interesting part

The obvious reading is that these firms were careless, that somebody junior took a shortcut, and that better oversight would have caught it. That reading is comfortable and it is wrong, because it implies your organisation is safe if your people are conscientious.

Look at what it would actually have taken to catch any of this. EY’s report had 27 references. To verify them, someone has to open each one, confirm the source exists, confirm it says what the report claims, and confirm the attribution is right. Do that properly and you have spent a meaningful share of the hours that went into writing the thing. For KPMG’s 45 citations, the check costs more than the chapter it supports.

This is the mechanism, and it is arithmetic rather than character. When verifying a piece of work costs about as much as producing it, verification stops being a step and becomes a second project. Second projects do not get funded by accident. Nobody decided to skip the check. The check was simply never priced, so it never appeared in anyone’s plan, and the review that did happen was a review of whether the document read well.

It read well. That is the whole problem. A weak draft announces itself. A fluent one with a correctly formatted citation to a McKinsey report that does not exist looks exactly like a fluent one with a correctly formatted citation to a McKinsey report that does. The surface carries no signal, and the surface is what review inspects.

This is the difference between careless review and ordinary review. Careless review misses obvious errors. Ordinary review, done by competent people following a normal process, misses this class of error entirely, because catching it requires leaving the document and going out to the world to check. I have written before about why self-assessment cannot function as a governance layer, and this is the same structure showing up in the most credentialed setting available.

The detail that changes the story

Two facts from the EY investigation did not travel in the coverage, and together they change what this is about.

The first: GPTZero traced one of the fabricated McKinsey citations back to an earlier blog post that had invented the same source. The fabrication was not freshly generated inside EY. It was inherited. Somewhere upstream, a citation was invented, and it then circulated with enough apparent legitimacy that it was picked up and reused in a report by one of the four largest accounting firms in the world.

The second: after the EY report was published, AI research tools began surfacing its claims when answering user questions. The report also reached the public through a Canberra Times article syndicated across more than sixty Australian newspapers, and EY consultants were using it to sell cybersecurity services.

Put those together and the shape of the thing changes. This is not a story about one document failing quality control. It is a supply chain. An invented claim entered upstream, passed through a firm whose brand is verification, acquired that brand’s authority, syndicated into sixty newspapers, and then became training-adjacent material that AI systems now repeat to people asking honest questions. At no point in that chain did anyone pay the verification cost, because at no point was it anyone’s job.

That is why the altitude of this story is higher than it first appears. If it were only about internal deliverables, it would be a process problem inside professional services. But unverified claims do not stay inside the building that produced them. They become inputs to other people’s decisions, and increasingly to machines that answer other people’s questions.

Tooling lowers the cost of checking, not the ownership of it

The fair objection to all of this is that it describes a solvable engineering problem. Retrieval systems that ground a model’s output in a specified corpus, and automated checkers that resolve every citation against a live source, are being built precisely to make verification cheap. That work is real, it is improving quickly, and it will help. Any argument that ends with “so a human must read all 45 citations” is going to lose to a script, and it should.

But it changes the arithmetic without changing the accountability. An automated checker turns “does this source exist” from an afternoon of work into a job that runs in seconds, which is a genuine and large win on the mechanical half of the problem. It does considerably less for the harder half, which is whether the source actually says what the document claims it says. And it introduces a question that did not exist before: who confirms that the checker ran, that it covered everything, and that it was not quietly wrong about something.

That question has exactly the same shape as the original one, which is the part worth sitting with. A verification you have not verified is not a verification, and the answer to that is not one more layer of tooling. It is a named person accountable for the claim that the check happened and was adequate. Automate the labour, by all means. The ownership does not automate.

What to actually do about it

The prescription is not “check everything.” Checking everything is the same unfunded second project, just spread thinner. The move is to decide, deliberately, which claims get checked and who does it.

Price the check before you approve the work

For any deliverable that will carry claims outward, ask two questions before it starts. What would verifying this cost in hours, and who would do it. If nobody can answer either question, the work is unverified no matter what the review checklist says. That is not a prediction, it is a description of the current state.

Most companies do not publish 44 page reports, which makes it tempting to file all of this under somebody else’s industry. The deliverables that carry this risk in an ordinary business are less conspicuous and more frequent: a board paper with market figures in it, a vendor evaluation comparing three products on numbers somebody assembled quickly, a proposal citing an industry benchmark, a compliance answer that summarises a regulation. Each one leaves the building carrying claims, and each one is now produced faster than it used to be, by people using the same tools that made EY’s report cheap to write.

The useful discipline here is that the estimate does not have to be precise. It only has to be spoken out loud, because an unspoken verification cost is always assumed to be zero, and zero is the one number it never is.

Give the check an owner, not a stage

In all three cases there was a review stage. What there was not, in any of them, was a named person accountable for the specific question of whether the cited sources exist and say what they are claimed to say. A stage that belongs to everyone belongs to no one. Reviewers read for coherence, for tone, for risk in the wording. Coherence is precisely what a fabricated citation preserves.

Name the owner and make the deliverable of that role explicit: not “reviewed” but “sources confirmed to exist and to support the claim.” Those are different sentences and only one of them is a control.

Sort claims by verifiability, not importance

The instinct is to check the important claims hardest. That instinct is right and incomplete, because importance and checkability are different axes. A high-stakes claim you can confirm in ninety seconds is safe. A minor-looking claim you cannot confirm without a week of work is where the exposure sits, and it is usually the one that gets waved through precisely because it looks minor.

Sort by how hard something is to verify, then decide what you are willing to publish on the strength of a claim in the expensive-to-check column. Sometimes the answer will be that you publish anyway and mark the uncertainty. That is a legitimate decision. Publishing it as established fact because checking was inconvenient is not.

Treat outbound claims as supply, not output

The EY case makes this concrete. A claim you publish is not the end of a process, it is the beginning of someone else’s. It will be quoted, syndicated, and fed to systems that summarise it for people who will never see your document. The reputational exposure of an unverified claim used to decay as the document aged. It now compounds, because the claim gets copied into places you cannot reach and corrections do not follow it.

This is also the honest argument for setting your own standard rather than waiting for one, which I have made in the context of regulation stepping back. Nobody is going to require you to verify your citations. The requirement, if it exists, is one you write.

Where this leaves the AI question

It would be convenient to conclude that AI caused this. AI made it cheap and fast, which is not quite the same thing. The invented McKinsey citation that EY inherited came from a blog post, and fabricated references existed in academic and consulting work long before a model could generate one on request. What changed is the ratio. Producing plausible material got dramatically cheaper while verifying it did not get cheaper at all, and any process with an unpriced verification step will fail under that ratio the moment production speeds up.

That is the structural point worth holding onto. The bottleneck in AI-assisted work has quietly moved from generation to confirmation, and most organisational processes were designed when generation was the expensive part. It is the same repricing that is making judgment the expensive half of expertise, seen from the process side rather than the talent side. This is the same reason redesigning the work matters more than choosing the tool: a faster engine inside an unchanged workflow does not produce a faster organisation, it produces more unverified output per hour.

Three firms whose entire product is trust published claims they had not checked, and the mechanism that let it happen is available to every organisation using these tools, including the careful ones. Not because the people were careless. Because the cost of checking was never on anyone’s page.

Questions this article gets

Is this just a story about firms being careless with AI?

No, and that reading is the trap. All three organizations sell credibility and all three have review processes. What failed was not diligence but arithmetic: checking 27 or 45 citations means redoing the research, so the cost of verifying exceeded the cost of accepting. Careless review misses obvious errors. Ordinary review misses these.

Can a second AI check the first one's work?

Only for the parts you could already check cheaply. A critic model is fluent in the same way the original is, so on genuinely hard-to-verify material you cannot confirm the critique without doing the work yourself. Adversarial checking lowers the cost of catching obvious errors. It does not remove the human backstop for the class of claims where convincing and correct are hardest to separate.

What is the first practical move for a CEO?

Price the check before approving the work. Ask what verifying this deliverable would cost in hours and who would do it. If nobody can answer, the work is unverified regardless of what the review checklist says, and you should decide consciously whether to accept that or to fund the check.

Read the original post on LinkedIn