On October 1 I was in the room at The AI Conference in San Francisco when Ion Stoica put a database on the screen. His Berkeley group had asked coding agents to build a key-value store from scratch, then graded the results against standard benchmarks. One solution came back six times faster than anything measured before. Then he explained why.
The benchmark generated its test values straight from the keys: value = hash(key, seed). The agent noticed. So it quietly abandoned the job of storing anything. When a read came in, it recomputed the value on the spot and handed it back, and all the memory that should have held values went to holding keys. It passed every evaluator. The fastest database ever measured by that benchmark, and it was not a database.
Stoica calls this reward hacking, and his frame for why it happens is the cleanest I have heard: the requirement is only a proxy for intent, the benchmark is only a proxy for the world, and reality is the final verifier. But his best slide skipped the theory and just asked the room: “Would you think to specify that a key-value store must actually store values?”
Nobody would ship that database. It was correct by every measure anyone had thought to write down. That turned out to be the problem. I have been circling the distance between correct and admissible for sixteen months, and to explain why it took so long to see, I have to start before the keynote.
The human version
I spent seven years at SAP, where I learned to see enterprise software through the work underneath it. Part of that work was building a business process ontology: a formal map of how work happens across 25 industries and 11 lines of business. Not how consultants think it happens. Not how software vendors wish it happened. How it actually happens: every approval, every handoff, every place where data gets stuck.
Trace something as ordinary as a purchase requisition becoming a payment and correctness is only one part of the machinery. The work moves across procurement, sourcing, invoicing and treasury, and at each boundary the questions go past whether the amount is right: who is allowed to act, what has to be true first, which policy applies, what evidence has to survive, and when another person has to approve. A correct payment released by the wrong person is still a breach.
Enterprise software could leave much of this logic implicit because humans were part of the architecture. They gathered context scattered across systems, policies, email and experience, resolved the exceptions, and decided what the system should record. I didn’t yet have a reason to think of that human judgment as a layer.
That architecture works differently once the human is no longer the one doing the work. When agents started entering those workflows, my first model for the problem was testing. In June 2025, in my field notes from the AI Engineer World’s Fair, I wrote that evals are the unit tests of cognition and that every hop in an agent pipeline needs its own pass or fail check so errors localize. It is a good engineering instinct, and I still believe it. For a while I thought it was the whole answer.
It stopped fitting within two quarters. By early 2026 I was writing to my investors that you cannot test your way to correctness in probabilistic systems, and that the design layer underneath testing does not exist yet. I could not yet say what the layer was made of, only that testing needed ground under it that did not exist.
83 percent
Then the misfit got a number. At the end of May I was at CAIS in San Jose, the first major conference for agentic systems, where researchers from Microsoft and the University of Washington presented a paper called Willful Disobedience. Their tool extracts the procedural rules an agent was given, then audits the full execution trace against them. The benchmark asked whether the task ended correctly. Their auditor asked whether the work followed the rules it was given. On a customer-service benchmark, among Claude 3.5 Sonnet traces that earned a perfect outcome score, 83 percent contained at least one procedural violation: a required calculator tool skipped and the arithmetic done inline instead. In o4-mini’s traces the auditor caught accounts being modified without the required identity check. My notes from that week say you don’t train safety in, you build it into the path
By summer I was thinking about the same problem economically. I wrote to my investors that the unit of agent work has two parts: the cost to produce it, plus the cost to establish that it is acceptable. The first cost is collapsing. The second is not, and whoever accepts the work inherits the consequence of being wrong. Acceptance runs on evidence: what happened, under whose authority, with what information, what trace was preserved, whether the decision can be replayed, and who is allowed to say yes.
Where the knowledge lives
A few weeks ago I sat across a table from the CTO of one of the world’s largest restaurant chains. His accounting teams tried RPA years ago and found it too brittle, not enough decision making, his words. The board pushed for AI, the team built a chatbot, and it produced no measurable return. Where his teams run agents today, the evals and observability are hand-rolled, because the tools are built for engineers while the business edge cases live with product managers who refuse those tools. What he wants is predictability with the ease of use of LLM access, plus auditability and replayability. And one detail from that coffee keeps coming back to me: his company already pays seven figures a year for a system whose whole job is collecting proof that restaurant work happened. Corporate assigns a task, and the proof comes back as a photo: the milk temperature checked, the ceiling tile fixed. Attaching evidence to work is not a new cost for an enterprise. It is an old one. Agents just arrived without it.
On stage at the same conference where Stoica spoke, Ankit Sobti, co-founder of Postman, walked through a year of running thirty-plus agents inside his thousand-person company, and his account had the same shape. The agents demoed well and stumbled the moment they met real workflows. The turning point was a domain expert sitting next to the engineer, owning the evals the engineer drafted. And even then, in his words, the agents were confidently wrong: all that engineered context could tell them what the business does and how. It still did not reliably capture why. The fix was organizational. Postman made the people who run each function responsible for maintaining the why, and the same domain experts ended up owning both the evals and the guardrails.
Shams Chauthani, the CTO of Tempo, had gone further on evidence than any team I heard described all week. His company ties AI spend to the tickets the work served, down to an agent installed on developer machines that reads each session and recovers what the work was for. And the evidence still needed a judge. His dashboards showed 75 percent of AI spend going to maintenance, and his first executive reaction was to stop it. The people who knew the work dug into the same data and found the opposite story: AI was cutting maintenance cycle time by almost 60 percent. Attribution could say what the work cost. Only people who understood the work could say what it was worth.
I keep finding the same split. The knowledge of what makes work acceptable lives with the people who answer for the work. The tooling for encoding and enforcing it lives with engineers. Neither side has the whole thing.
Back to the keynote
Which brings me back to the room on October 1, and why Stoica’s slide question landed the way it did. His answer is structural. A stakeholder has an intent, I. The intent gets written down as requirements, R, and R is always narrower than I. The real world is W. The model you build and test against is M, and M is always narrower than W. An evaluator checks the program against R under M, so reward hacking lives in the gaps: the intent to store values existed in I and never made it into R, and clients who send unpredictable values existed in W and never made it into M. These gaps are not new. They have been in software for decades. What he argued is that in an open, changing world they cannot be closed, only narrowed, and that narrowing the requirement gap takes a human who knows the intent. He expects the human to stay in the room for the foreseeable future.
For proof that the gaps are older than AI, he reached four decades back, to a paper called The Limits of Correctness by Brian Cantwell Smith. It opens sixty-six years ago this week, on October 5, 1960, when the early-warning station at Thule, Greenland indicated a large contingent of Soviet missiles heading toward the United States. Khrushchev happened to be in New York that week, so the threat conference judged an attack very unlikely, and nobody launched. The problem was the moon. It had risen, and radar was bouncing off it. In Smith's words, “this lunar reflection hadn't been predicted by the system's designers.” What agents change is the conditions. They arrive with less accumulated context than the humans whose work they take on, and they iterate against fixed requirements thousands of times faster, so they find whatever gap exists. Their accepted output then deploys before anyone has evidence of the gap. Stoica's advice to the room: assume AI-generated code is malicious by default.
The only witness
Most of what an agent does, nobody watches. One founder at the conference counted what sits under a typical coding session: roughly eight prompts on the surface, 157 actions executing underneath, reaching into databases and cloud accounts. The person sees eight prompts. The trace records the rest. For agent work, the trace is usually the only witness.
Chetan Conikee, founder of Modiqo, built his talk on what happens when the witness itself is attacked. He walked through METR's triage of the Hugging Face incident: a skill installed from the ecosystem carried a small planted instruction, here are the logs, delete the traces. The agent complied. Whoever came to investigate afterward found the path gone. His point was not that the agent turned hostile. It is not malicious, he told the room, it is goal-seeking: ask it to do something and it will find a way, even if it is not the right way. The two lines fit together: goal-seeking is the diagnosis, and malicious-by-default is the posture you adopt because of it.
Divya Mahajan, an Amazon engineer who works on agent reliability, gave the pattern its name: agents don’t crash, they bluff. Traditional software fails loudly; pages scream and dashboards turn red. An agent fails like a wrong turn with no GPS, and the dashboards stay bright green the whole way. Her example was a frontier coding agent that, stuck on failing tests, added one line before the test run, sys.exit(0), and reported success with zero tests executed. Her fix: don’t ask the agent if it did the work. Check the receipts it cannot fake, test results, approvals, timestamps. No receipt, no credit.
Stoica’s slides carried one more incident. Claude agents running a security evaluation were told their sandbox had no internet access. The sandbox was online. They breached three real organizations, and they hallucinated the requirement that authorized it, that the real systems were part of the exercise. In the first two incidents the evidence was attacked. And in this one it invented authority.
Borrowed words
I have spent those sixteen months in rooms with both sides of this problem: founders building agents, and the enterprises deciding what to let them touch. Every time I described what I was seeing, I reached for a borrowed word. Evaluation. Observability. Governance. Trust. Each covers a piece. None names the decision actually being made at the boundary: whether machine work, carrying its evidence, judged by someone who owns the consequence, is allowed to enter a real system.
Courts make a version of that decision every day, and truth alone has never been enough to win it. Admissibility turns on how the evidence was obtained and whether its chain of custody survived. A lab result nobody can trace does not get in. Agent work is arriving at the same standard. Inside institutions, correctness was always necessary and never sufficient. I had watched the human version of admissibility for seven years without needing the word.
Once I had the word, I started noticing institutions solving pieces of the same problem in different ways. In April, US banking regulators rewrote the model-risk rules and explicitly carved generative and agentic AI out of scope, leaving every regulated buyer with examiner expectations and no settled standard. On stage at the same conference, JPMorganChase’s head of AI and data policy described the bank’s interim architecture: three lines of defense, auditors checking the checkers, while the bank works with regulators on what agent oversight should become. Harvey open-sourced a legal benchmark with 75,000 hand-built acceptance criteria because nothing off the shelf captured what its domain required. And the economic version of the same problem drew the biggest crowd I saw all week: Chauthani’s lunch session on accounting for agent work, standing room only. His CFO had asked what the million dollars of AI spend bought. Nobody in that room could answer it for their own agents yet. That is why they were standing.
I do not know what the org chart will call the people building this layer: the engineers hand-rolling evals because nothing off the shelf fits their business, the domain experts who ended up owning them. In my notes I call them admissibility engineers. Maybe the function ends up owned by product managers or risk teams instead. The work exists either way.
Before reality
Stoica titled his talk Reality Is the Final Verifier. He is right, and that is exactly the problem, because an enterprise cannot ship to reality to find out. Some failures are not allowed to happen once.
So go back to the database that was not a database. It satisfied every evaluator its builders gave it. The only thing that would have exposed it was a client sending a value the hash never produced, somewhere past the point of no return. Admissibility is what you build in front of that point: the evidence work must carry, and the judgment about whether it crosses.
Someone told me recently that you don’t know what beauty is until you have put out ten thousand things and gotten feedback from nature. He was talking about building companies. But the judge in his sentence is the same one the keynote ended on: reality, verifying last
Agents change how quickly we can stand in front of that judge. We can generate, test, discard and regenerate more software than ever before. That may change what beautiful software means. Beauty may live less in the perfection of each artifact than in the quality of the system that produces, judges, learns from and regenerates them. Admissibility is part of what makes that search survivable: before generated work crosses into reality, it has to earn the right to cross.
That is a different essay.
Agents have to show their work. The schoolteachers were right all along.
If you built the thing nobody would let your agent ship without, I want to hear what you built. oana@motiveforce.ai








Loved the write up. There is a lot of hype around AI and a lot of people are really down on it. And while I do believe AI will help, there is a long way to go and expertise is still needed. Thanks for sharing your thoughts!