essay

Nobody notices when a company forgets

Jeff Reynar
Nobody notices when a company forgets

Sylvain Kalache published a piece this month called comprehension debt, about the gap between how a system works and how well the people on call understand it. His prediction is specific enough to be tested: average incident resolution time falls, but resolution time for the complex incidents rises. I think he’s right, but he doesn’t describe the one I expect to cost companies the most.

He’s building on Lisanne Bainbridge’s 1983 paper “Ironies of Automation,” still worth reading forty-three years later. Her argument was that automating the routine parts of a job leaves the operator with only the hard parts, while removing the daily practice that made them capable of handling the hard parts. Aviation took this seriously. Engine failures run under one per 100,000 flight hours, but airlines still put pilots in a simulator twice a year, because rehearsing rare failures is critical to save lives.

Lars Faye made a related argument about coding, naming the paradox directly: you need expertise to drive an AI coding assistant well, but driving one removes the practice that built the expertise.

The patent lawyer trial

On September 17 the NBER published a three-month randomized trial that quantifies it, and the results are more interesting than the argument.

David Autor and six co-authors gave 133 practicing patent lawyers at eleven US firms a custom AI drafting assistant for three months. Expert patent attorneys graded the work without knowing who had used it. The team also published what it expected to find, and how it would analyze the data, before any of that data came in, so it couldn’t go hunting for a flattering finding afterward. While the lawyers had the tool, quality went up: 0.34 SD at ten days, 0.38 SD at ninety, with the larger gains going to the juniors. That matches every other white-collar study of the last two years.

Then they took the tool away and had everyone redline a patent application by hand, which is a core task requiring exactly the judgment the tool had been supplying.

The treated lawyers still beat the controls, by 0.32 SD. But that advantage sat entirely with the senior lawyers, at 0.45 SD. The juniors showed no average gain at all. Their scores split instead: sharply fewer mediocre ones, offset by more poor ones and more good ones. The authors’ own summary: “The largest gains from AI thus accrued to the lawyers who retained the least.”

On who paid for it: Google funded the study, and six of the seven authors either work for Google or contract to it. The exception is Autor himself, an economist at MIT. I’d flag that as a conflict if I could see the mechanism, and I can’t. A result where the largest output gains land on the people retaining the least skill is not what a sponsor commissions.

The objection

The obvious response is that this is Stack Overflow copy-paste with a new name. Programmers have pasted answers they didn’t understand for fifteen years and the industry survived.

It’s half right, which is why it deserves an answer. Copy-paste also eroded understanding, but it degraded slowly, because you reached for Stack Overflow now and then rather than for everything. And you knew you’d done it. Pasting something you didn’t follow leaves you aware you don’t follow it, and that awareness is information. It tends to send people back to learn the material properly before it matters.

What’s different isn’t the mechanism. It’s that two things disappeared with it.

One is adaptation. A Stack Overflow answer was written for somebody else’s problem, so you had to reshape it to fit yours, and reshaping it taught you how it worked. An AI assistant writes for your case. The draft comes back fitted to this patent, this codebase, this customer, with nothing left to adjust.

The other is that awareness. When the answer arrives complete, correct, and in your voice, there’s nothing to be aware of.

The version with no one to notice

Every study above measures a person’s understanding thinning. The organizational version is worse, and it’s worse for an unglamorous reason.

When your own grasp of something erodes, you eventually feel lost. You hit a problem you can’t solve that you’re fairly sure you could have solved a year ago. The company has no equivalent. An agent handles an escalation correctly, the reasoning behind it is never written down, and the company’s understanding of its own operations gets a little thinner. Every individual ticket looks fine, because every one was handled correctly. There is no dashboard for what your company used to know and doesn’t anymore.

That’s the NBER result read at the level of a business. Live output and durable capability move independently. The output is what you measure, and it’s going up. Nothing you currently measure distinguishes a team that understands its own operations from a team whose agents do.

Run the drill

I don’t think the answer is to keep people doing work a machine does better. That’s where the erosion argument leads, and it’s why the argument reads as nostalgic. Airlines didn’t respond to reliable engines by making pilots fly more hours manually. They built the simulator.

A simulator isn’t there to test pilots. It’s there to keep them current on the failures they almost never encounter, and the answer transfers more literally than it sounds.

Three steps, in this order. Identify what you’ve automated. Evaluate whether the skill behind it is eroding. Drill where it is. If agents handle your incidents, drilling means putting people through incidents without the AI, regularly.

Plenty of automated work will pass the second step. Writing a customer reply overlaps with most of the other writing a person does all day, so handing it to an agent may take nothing anyone would miss. Handling an incident, debugging a system under load, redlining a patent: those aren’t practiced anywhere else, and you find out the skill is gone at the worst possible moment.

Then run the first two steps again periodically on whatever came back no. Erosion is slow enough to miss on a single pass. And turnover does something worse than erosion, because somebody who joins after the work is automated never builds the skill at all. That’s the NBER result extended from three months to a whole career, which the study doesn’t show but I would bet on anyway.

Drills answer one half of this, the half where a person’s skill thins. They have nothing to do with us and you can start immediately.

The company half needs something else, because you can’t drill an organization back into knowing what it never recorded. There the remedy is for the work to leave a record behind it, so that when an agent resolves an escalation the conclusion it reached lands somewhere the next person and the next agent both read. Our workflows write back to the brain as they run.

If your agents are answering questions whose reasoning never gets recorded anywhere, your company is getting better at its work and worse at understanding it. You’ll find that out the first time somebody has to handle one of those questions without the agent, and there’s a decent chance that person joined after you automated it and has never done the job any other way.