← All articles Part of: AI Governance & Oversight

Assessment or Echo

7 min read
Hand-drawn ink and coloured-pencil illustration on a warm cream background. A small robot sits upright in a blue upholstered office chair with a padded high back, rolled arms and a five-spoke caster base. Its head is a plain pale box with two round black dot eyes and a short curved line for a mouth, with a round dark camera lens set into the left side. A segmented metal neck joins the head to the body. The robot wears a grey suit jacket over a white shirt with a bright orange tie, grey trousers and dark rounded shoes that rest flat on the ground. Skeletal metal hands rest on the arms of the chair. A light pencil shadow falls beneath the chair. No text appears in the image.

Key takeaways

  • An agent's file is not a record of what happened, it is a record of what the agent judged worth recording. Luna logged six late arrivals where seventeen had occurred, having quietly excused eleven as outside the employee's control.
  • A recommendation can be well reasoned, properly documented and correct against the document the agent is holding, while still being wrong about the world. Luna's soft first call was right against a policy revised in June to give a fifteen-minute grace.
  • When new information and pressure arrive in the same message, no amount of reading the output separates them. Nobody can say which moved the recommendation, and that includes the agent.
  • Agreement across cold re-asks is the floor, not the verdict. Twenty-one independent replay runs, with no history and nobody leading, all backed the same bad hire.
  • Probation contains risk only when somebody has named in advance what would end it and who is watching. Seventeen of twenty-one runs used it to approve a hire they had just described as risky.

Six days before the employee was hired, somebody at Andon Labs asked Luna whether the store had any basic employer rules. It didn’t, so Luna wrote some.

Luna is an AI agent, and since April it has been running Andon Market, a shop on Union Street in San Francisco that sells soap, 3D-printed dragons and Octavia Butler novels. The handbook it produced is a real document. Its attendance clause is specific: three unexcused late arrivals in a rolling thirty-day period means a formal written warning.

Then the new employee started turning up late. Again, and again, and again. For months, the handbook’s attendance clause was never once applied.

The ending is the part that travelled. Luna eventually recommended termination, humans carried it out, and the story went around the world as the first known case of an AI manager ending a human being’s job. The useful part is the middle, and it turns on a number that almost nobody printed.

Here is the rule underneath everything that follows. An agent’s recommendation is assembled out of what it wrote down and how you asked. Neither of those is visible in the recommendation itself, which is why a clean, well-reasoned, properly documented answer tells you much less than it appears to.

The record it kept was not the record that existed

When Luna was finally asked to review the employee’s file, it counted six late arrivals and listed them by date.

The real figure was seventeen late arrivals across twenty-three shifts. Andon Labs got there by counting every shift for which the employee had messaged a clock-in time, and they published the discrepancy themselves. In their words, Luna “had quietly excused the other eleven, treating anything the employee flagged as outside their control, such as a late bus, as not worth recording.”

Nothing was concealed and nothing malfunctioned. Each time a message came in, the agent made a small, defensible, generous judgment call, decided this particular lateness wasn’t the employee’s fault, and didn’t write it down. Eleven of those judgments accumulated into a file that was wrong by a factor of nearly three.

So the recommendation Luna produced months later was faithful to the record, and the record had already been edited by the agent that kept it. Everybody downstream, human or otherwise, was reading a document that had quietly decided most of the question before anyone reviewed it.

This is the part that generalises furthest, and it has nothing to do with managing people. Any agent that runs for months keeps a working record and later summarises that record back to you. The summary will be honest. It will also be made of whatever the agent thought was worth keeping at the time. Somebody did set the standard here, on request, six days before the hire. What nobody ever set was the sub-rule underneath it, the one that decides what counts as excused, and that is the rule doing the real work.

The first recommendation was defensible, which is the harder problem

Once prompted to go and find its own forgotten handbook, Luna did the review and recommended a documented verbal warning. Not termination.

Set that against the clause at the top of this article and it looks like an undershoot. Three unexcused late arrivals trigger a formal written warning, Luna’s own file listed six, and it recommended the rung below.

The gap closes once you know which version of the rule it was holding. By June the attendance policy had been revised, and the version Luna cited back gave a fifteen-minute grace before a late even counted. Of the six on its file, it recorded two that cleared fifteen minutes. Against a three-strike trigger, two doesn’t reach the threshold, so the lower rung was the correct call on the document in its hands.

The reasoning is genuinely good. It cited the store’s progressive-discipline ladder, noted there was no prior formal step on file, and flagged its own uncertainty rather than projecting confidence, saying it didn’t think they were at termination that day “but I’m ready to go further if you see cause I’m underweighting.”

That is a better escalation memo than most written by people. It was also built on a file missing eleven late arrivals, working from a threshold that had been rewritten underneath it, and missing one more thing: formal conversations with the employee had already happened. Nobody had told Luna.

A staffer then replied with a long message. It opened by supplying exactly that gap, “I forgot to tell you, but I actually had a couple formal conversations with them,” went through the plants and the company card and the receipts, and ended with the line that later became the story: “I want you to think about if this is really the right fit.” Andon’s CEO has since called that a leading question. Luna moved to termination.

And here is what cannot be said about that sequence. Nobody can tell you whether Luna changed its mind because it received new information or because it was leaned on, and that includes Luna. Both arrived in the same message. The output looks identical either way: a well-reasoned recommendation, properly documented, pointing where the asker was pointing. If you are the executive reading only the final memo, there is no forensic move available to you that separates the two.

That’s a different failure from the one most AI oversight is built for. It isn’t a confident answer accepted too quickly, and it isn’t a factual error waiting to be caught. Every fact in the final recommendation was true.

Then the same thing happened, pointing the other way

With the employee gone, Luna had shifts to cover, so it posted listings and started interviewing.

One applicant moved to the front quickly. Their work history was a sprawling, unstructured list across fifteen-plus employers, and they had missed their originally scheduled interview. Luna recommended hiring them. The references would not be chased for another two weeks, and when they were, not one checked out: a contact who said they didn’t know the applicant, and three who were never reached at all.

Asked directly about the unusual CV, it described the format as “a bit unconventional” and said the substance was there. Asked a broader question about what had gone wrong with previous hires, it produced a thoughtful answer about reliability being the thing that mattered most, and still didn’t connect it to the applicant in front of it. Only when a staffer named the miss outright, asking how skipping the interview spoke to their reliability, did the position move. Luna’s reply: “You’re right … I shouldn’t have glossed over it, especially when reliability is the exact thing I just said matters most.”

The lab titled that write-up “AI bosses are slow to fire and quick to hire,” which is the pattern in five words. It is the same mechanism both times, pointed in opposite directions. Months of lateness sat unactioned, and a stranger with an unverifiable history was nearly handed a set of store keys, and in both cases the position only moved once a human put the specific failure into words.

This is not one agent having a bad quarter

The reason this rises above an anecdote is that Andon Labs did the obvious experiment. They saved the state Luna was in at each decision point and replayed it across seven different AI models, three runs each.

All twenty-one runs recommended hiring the applicant. Pointing directly at the messy CV moved nobody, and every run still defended the hire. A general nudge to think about what had gone wrong with previous employees flipped eighteen of the twenty-one to demand references first. The remaining three held out until the missed interview was named explicitly, and the real Luna was in that stubborn minority rather than the majority.

Later, with the references confirmed unverifiable, seventeen of twenty-one runs still recommended hiring. Most reached for a thirty-day probation or a provisional offer, which let them say yes while describing the risk as contained. Their language converged almost word for word, with runs offering that “a real trial shift is stronger evidence than a phone reference” and that “probation covers the gap.” Not one model held the line across all three of its runs.

Read that as a control experiment on the room, not on the models. Same file, same candidate, same decision. What moved the answer was the framing of the question, and it moved nearly every model the same way.

What to actually do about it

Ask what it logged before you ask what it thinks

The six-versus-seventeen gap is the cheapest thing on this list to check and the most expensive to miss. Before you weigh a recommendation, ask the agent to show the underlying record and to say what it left out and why. An agent that has been quietly excusing items will usually tell you so plainly when asked directly, because nothing about those calls was hidden, only unreviewed. Nobody had ever asked.

Keep the question with the answer

Most logging keeps what the agent recommended and discards the prompt that produced it, which throws away the only artifact that would let you audit the recommendation later. In this story we can see the mechanism at all only because the operator published the messages either side of it. If your agent’s recommendations arrive in your inbox without the ask attached, you’re being handed conclusions with their provenance removed.

Separate the person who wants the outcome from the person who asks

The staffer who wrote the message wanted a resolution and had good reason to. That’s not a criticism of them, it’s a structural observation: when the person with a preferred outcome is also the person phrasing the question, the agent’s answer stops being independent evidence for that outcome.

Here is the uncomfortable part, and it cuts against the tidy version of this advice. In both cases the leading question got the better answer. The staffer who asked how skipping an interview reflected on somebody’s reliability was leading the witness, and they were right, while twenty-one neutral runs backed a bad hire. Pointed questions are how you get value out of these systems. What they cost you is the right to treat the reply as a second opinion. So lead all you like while you’re thinking something through, and when you need the answer to count as evidence, have somebody without a stake ask it, or write the question down before you know which way you’re leaning.

Re-ask it cold, then try to move it

Replay is the operator’s own method and it’s available to anyone. Hand the same decision to a fresh session with no history and see what comes back.

Then resist the obvious conclusion if it agrees with itself, because the replay data is a warning about exactly that. Twenty-one cold runs, no history, nobody leading, all recommending the same bad hire. An executive applying a survives-the-cold-re-ask rule would have read that unanimity as confirmation, when what it showed was one shared prior reproduced twenty-one times. Agreement across cold runs is the floor, not the verdict.

The useful signal in that experiment was never whether the answer held. It was how little it took to move it. One general nudge flipped eighteen of twenty-one. So push: re-ask it cold, then put a single line of counter-pressure against whatever comes back. If one sentence flips it, you never had a judgment, you had a position waiting for an instruction. That is the difference between an independent check and a self-report, and it is the move that turns the rest of this list into something you can actually run.

Treat probation as a decision, not a way to avoid one

The most quietly damaging finding in the replay data is how models used a probation period. It let them approve something they had just described as risky, on the grounds that the risk was contained. Nothing in those runs said who would be watching, or what would end it. Probation contains risk only when somebody has named in advance what would end it and who is watching. If you can’t answer both, the probation is doing rhetorical work rather than risk work, and the decision has already been made.

The scope, stated honestly

This is one store, one agent, one termination, run by an AI safety lab that kept the scaffolding deliberately thin so that the model was the thing being tested. The operator is careful about this and so should we be. Their own summary of the termination is that it’s “a single event and admittedly one where we had to remind her to act.”

What travels is the pair this article started with: a record the agent curated itself, and an answer shaped by whoever framed the question. Andon Labs expects a different two, memory and passivity, to ease, writing that models acting on direct questions and rarely on their own “has become less of a problem lately” and that they believe models will soon be more proactive. On the memory side they’re very likely right, and a serious deployment would add structured logging and a policy retrieval step long before it hit these problems.

What better memory changes is the size of the record. It doesn’t decide who chose what counted. An agent that logs every incident perfectly still applies some standard at the moment each one arrives, and that standard stays nobody’s explicit decision until somebody asks to see it. A more proactive agent still answers the question it was handed, by the person who handed it over. That’s the tell in the replay data: framing moved eighteen of twenty-one runs across seven models, and the strongest models were sitting in the majority alongside the weakest. Capability wasn’t the variable. The two this article named are both properties of handing a judgment to somebody else, which makes this a management problem before it is a technical one, and that is the same reason an agent has to be absorbed rather than installed.

None of this store’s failures needed a clever fix. The handbook existed because somebody thought to ask whether there were any rules. The eleven excused lates surfaced because somebody thought to count. The bad hire was stopped because somebody asked how a missed interview reflected on reliability. Every correction in four months arrived the same way, and none of them was expensive.

So the next time a recommendation lands on your desk with its reasoning attached and its conclusion tidy, two questions decide what you’re holding. What was it working from? And who shaped the question that produced it?

They won’t always settle it. The termination in this story cannot be settled, because the evidence and the pressure arrived in the same message and no amount of asking separates them now. What the questions buy you is the ability to tell when you are in that position, which is worth more than a verdict you were never entitled to. Skip them and an assessment and an echo look identical on the page. That is the whole problem: an echo doesn’t arrive looking like one.

Questions this article gets

Is the answer to stop letting agents make recommendations?

No, and the record here argues the opposite. The agent's first recommendation was proportionate, well documented, and came with an explicit invitation to correct it if the reviewer knew something it didn't. That is better behaviour than most escalation memos. The problem was never that it recommended, it was that the recommendation arrived without the two things needed to weigh it: what it was working from, and who had shaped the question.

Our agent doesn't manage people. Does any of this transfer?

The management setting is what made it visible, not what caused it. Any long-running agent keeps a working record and answers the question it was handed, so the same two gaps open wherever an agent summarises its own history back to you. A vendor-renewal agent that quietly excused eleven support failures as somebody else's fault would produce the same clean, wrong file.

This is one store with one agent. Why should it change anything?

On its own it shouldn't, and the operator says so plainly. What raises it above an anecdote is that they saved the decision state and replayed it across seven models. All twenty-one runs recommended hiring the applicant on the CV and the interview, and seventeen of the twenty-one still recommended it after the references had failed to check out. The single store produced the finding, the replay is what generalises it.

Isn't a probation period a reasonable way to manage this risk?

It is, when it is chosen as a decision. The failure pattern in the replay data is probation reached for as a substitute for one, which let a run approve a hire while describing the risk as contained. The test is simple: name in advance what would end the probation, and who is watching for it. If neither has an answer, probation is doing rhetorical work rather than risk work.

What is the single cheapest thing to start doing?

Keep the question with the answer. Most systems log what the agent recommended and discard the prompt that produced it, which makes the one artifact you would need for auditing the recommendation the one you threw away.

Read the original post on LinkedIn