Key takeaways
- An attorney facing sanctions asked ChatGPT directly whether the case it had invented was real. It told him the fake case "does indeed exist" and was available on Westlaw and LexisNexis. Asking a system to confirm itself is not a second opinion, it generates its answer the same way it generated the error.
- A signature was already required on the filing, a human checkpoint built into the rule. The attorney who signed it did not review the cases first. A checkpoint that exists on paper without someone actually exercising judgment there is not a control.
- Checking every AI output fails for two separate reasons: it is rarely funded once it costs close to what producing the work costs, and it cannot keep pace once a process compounds faster than a person can read each result.
- Both EU law and US banking regulation already scale human oversight by consequence rather than applying it uniformly, and for one category it singles out, identifying a person from biometric data, EU law requires two independent people, not one, before anyone acts on it.
- A gate is a specific moment, not a calendar date: an assistant that stops needing review before it sends, a spending limit that goes up, a support bot that starts closing tickets instead of drafting replies. That is where the sign-off belongs, not on each output once it ships.
In this article
In March 2023, an attorney named Steven Schwartz filed a legal brief in federal court that cited court decisions supporting his client’s case. Opposing counsel could not find any of them. Neither could the judge. Schwartz went back to the source that had produced the citations and asked it directly: was one of these cases, Varghese, actually real.
ChatGPT told him it was. The case “does indeed exist,” it said, and he could find it on Westlaw and LexisNexis. He later testified he was already suspicious by then. He asked anyway, and the system answered with the same fluent confidence it had used to invent the case in the first place. Judge P. Kevin Castel’s opinion in the case, Mata v. Avianca, published in June 2023, records both the fabrication and the attempted check. Schwartz and his colleague were sanctioned $5,000, jointly and severally with their firm.
That exchange is worth sitting with before anything else, because it is the whole failure in miniature. Schwartz did not skip verification. He ran one. He simply ran it against the thing that had produced the error, and asked it to grade itself. Fixing that was never going to be a matter of asking harder. It’s a matter of asking somewhere else entirely, which is what the rest of this piece is about.
Why the check didn’t check anything
A language model does not have a separate faculty for telling truth from invention. It generates an answer to “is this case real” the same way it generated the fake case to begin with: by producing text that reads as plausible and confident. Asking it to confirm itself is not a second opinion. It is the first opinion, restated with more certainty than the first time.
This scales far beyond one lawyer’s citation check. Anthropic has reported that, as of May 2026, more than 80% of the code inside its own systems was written by its own AI model. At that volume, no person is reading every line before it ships, so checking that code increasingly falls to other automated systems, the same shift in kind as Schwartz handing his own check back to ChatGPT. Anthropic’s own account of where this leads, published as “When AI Builds Itself,” calls it recursive self-improvement, an AI system designing its own successor rather than waiting for engineers to build it, and admits plainly that “it’s possible that we can’t build, integrate, and verify the tools that we’d need to understand which trendline we are actually on.”
A system grading its own output is a feedback loop rather than a governance layer, whether or not the system is honest, a distinction worth citing rather than re-arguing here. What is less settled is what happens once you swap in a second, genuinely separate AI system to do the checking. It helps, but it doesn’t close the loop. A second model is still generating an answer, not consulting an independent record, and unless something outside the model family confirms what it reports, you have relocated the feedback loop, not closed it.
The other obvious answer, check everything by hand, fails for a different reason. Verifying a piece of work costs close to what producing it costs once the material is dense enough, and when a check costs that much, it quietly stops happening, no matter who is nominally responsible for it. That rules out universal review on cost alone. Speed is the second, separate reason it fails, and it’s the one the rest of this piece is actually about.
Two industries that already solved a version of this
Regulators covering high-stakes automated decisions have already worked through a version of this exact problem, in two different domains, and neither one landed on “review everything.”
The EU AI Act’s Article 14, covering human oversight of high-risk AI systems, does not ask for uniform review. It requires that oversight be “commensurate with the risks, level of autonomy and context of use” of the system in question. It specifies what a person overseeing such a system has to actually be able to do: understand what the system can and cannot really do, and resist what the law names directly as automation bias, the pull to simply trust the output. The person also has to be able to correctly interpret what it produced, and to override, disregard, or stop it. For one category the law singles out for an extra safeguard, identifying a person from biometric data, it does not settle for one reviewer. It requires that the identification be “separately verified and confirmed by at least two natural persons” before anyone acts on it.
US banking regulation reaches a similar design from a completely different direction. The Federal Reserve and the OCC’s 2011 guidance on model risk, SR Letter 11-7, calls for “effective challenge” of any model used in a bank’s decisions: “critical analysis by objective, informed parties who can identify model limitations and assumptions,” kept separate from whoever built the model. It says outright that this should not be uniform either. “Where models and model output have a material impact on business decisions… a bank’s model risk management framework should be more extensive and rigorous,” while a model with limited reach gets a lighter check.
Neither document was written with an AI system checking another AI system in mind, and neither is offered here as a template to copy line for line. What both share is the design decision worth stealing: the check does not sit on every output. It sits at a small number of specific, well-chosen points, and the strength of the check scales with what happens if it’s wrong.
Four moves, in order
Map where AI touches a decision, an output, or a system, and ask, process by process: who confirms this today, and against what? Not “reviewed.” Confirmed against something the system itself did not generate: a signed document, a lab result, an independent record, a customer’s own account of events. Where checking has never been priced, the honest answer is usually nobody, and it’s worth saying that plainly before doing anything else.
Sort what you find by one question: if this is wrong and nobody catches it for a month, what does that cost, and can it be undone? A mistake you can catch and undo cheaply is fine running on the system’s own report of itself. A mistake that compounds, that gets acted on before anyone looks again, or that cannot be undone once it ships, is the category everything below applies to.
For anything in that second category, keep asking who checks the checker until the answer is a person or an artifact a person is accountable for, not another AI system pointing back at itself. If the answer comes back “another AI system,” treat that as a deferral, not a resolution, exactly the deferral Schwartz made when he asked ChatGPT to confirm ChatGPT. The chain has to end somewhere a human would put their name on the result, the way a court filing requires an attorney’s own signature, or SR 11-7 requires a named validator distinct from whoever built the model. A person at the end of the chain isn’t automatically a real check either. What makes a checkpoint real is the same thing SR 11-7 requires of its validators: the incentive, the competence, and the standing to say the result is wrong, not just a name attached to it.
Move the sign-off to the gate, not the output. A gate is a specific moment, not a calendar date: the day a drafting assistant stops needing review before it sends, the day an agent’s spending limit goes up, the day a support bot starts closing tickets instead of only suggesting replies. Each of those is a real increase in what the system can do unsupervised, and each is exactly where the sign-off belongs, because once a process compounds faster than a person can read each result, reviewing every output stops being expensive and becomes arithmetically impossible. The question changes from “is this particular answer right” to “have we earned the right to let this run with less supervision than it had yesterday,” and that question can still be answered at human speed even when the output itself can’t be. This doesn’t say how often to ask it, and no fixed cadence would survive contact with a real organization anyway. What it does say is that the moment worth watching for is any expansion of scope, autonomy, or speed, however small, because that’s how supervision quietly erodes without anyone deciding it should.
The gate still has to be manned
It’s worth naming how this fails even when a gate already exists on paper. Peter LoDuca, the attorney of record in the Avianca case, was required to sign the fabricated filing under penalty of perjury, exactly the human checkpoint the court’s own rules assume is there. Judge Castel’s opinion states plainly that before he signed it, LoDuca “did not review any judicial authorities cited in his affirmation,” and made no inquiry into how his colleague had researched them. The gate existed. Nobody stood at it. The $5,000 sanction landed on the whole chain, the firm and both attorneys jointly, not just on whoever happened to be typing, because, as the opinion put it, existing rules “impose a gatekeeping role on attorneys to ensure the accuracy of their filings.” A signature with no check behind it is a formality standing where a gate should be.
Whether the AI touching your own operation is honest is unfalsifiable from where you’re standing, and chasing an answer to it is a dead end. What you can actually find out is where your own chain ends today: in a person who checked something real, or in a system reassuring you, confidently, that its own work is fine. Find that point for the process where a wrong answer would cost you something real. If nobody is standing there, that isn’t a detail to get to later. It’s the gap this whole exercise exists to find. And when the thing being checked is an agent acting on live systems, the gate has to sit outside it, with a named person who answers when it fires.
Questions this article gets
Doesn't using a second AI system to check the first one solve this?
Only partway. A second model is still generating its verdict the same way it generates everything else, by producing plausible text, not by consulting an independent record. It can catch categories of error the first model has no incentive to hide, which is real and worth having. It cannot certify its own honesty, for the same structural reason the first model couldn't. The chain still needs to end somewhere outside the model family: a person, or an artifact a person is accountable for.
Isn't reviewing every AI output the safest approach?
It sounds safest and it is usually the first thing to rule out. Verifying a piece of work at anywhere near the cost of producing it turns the check into an unfunded second project, and unfunded checks quietly stop happening, whoever is nominally assigned to them. It also does not survive contact with speed: once a process compounds, output arrives faster than a person can read it, and no amount of goodwill fixes an arithmetic problem. The fix is choosing a smaller number of well-placed checkpoints, not reviewing everything thinly.
What is the one thing to check first in my own operation?
Pick the AI-touched process where a wrong answer would actually cost you something real, not the one that feels most futuristic. Ask who currently confirms its output, and against what. If the honest answer is nobody, or another AI system with nothing behind it, that is the gap. Fix the highest-consequence one first, and place the check at the point where the system's autonomy or scope is about to expand, not after each output ships.