Agents didn't stall. We gave them the wrong test.
Mark Zuckerberg told Meta employees this month that AI agents haven’t progressed as quickly as he’d hoped. The memo made it to TechCrunch, and the reaction split along familiar lines. Skeptics read it as proof the agent era is stalling. Believers read it as another story about an impatient executive with unrealistic expectations. I think both readings miss what the evidence actually says.
The numbers behind the disappointment
Two data points anchor the pessimist case.
The first is capability. Scale AI measured frontier agents against complex multi-day tasks, real freelance projects sourced from online platforms, and found they complete about 4% of them at a level matching the human gold standard. Whatever the demos suggest, an agent you brief on Monday and grade on Thursday will usually disappoint you.
The second is deployment. Writer’s enterprise survey found that 97% of executives say their company deployed AI agents in the past year, while 79% of organizations report that adoption is a struggle.
Put the two numbers side by side and you get the pessimist’s case. Nearly every company deployed agents, and agents fail at 96% of substantial projects. That distance between adoption and delivery is the stall Zuckerberg is naming.
Look deeper, though, and the deployment data partly vindicates the agents. The deployments that succeed focus on small, bounded jobs: notifications, reminders, routing, triage. Agents may not be good enough to meet Zuckerberg’s expectations, but they are delivering results. The question worth asking is why they succeed there and fail the big test.
A test no manager would run
That story rests on the assumption that the briefed multi-day project is the right test of progress.
Look at the conditions of that test. Scale AI’s tasks are real freelance briefs, delivered cold. The brief is all the context the agent gets. There’s no client to ask a clarifying question and no colleague one desk over. And the agent runs for days between checkpoints, compounding every early misread without correction.
We evaluate agents, in other words, in a way we’d never evaluate a new hire. We give them a project with minimal context, no opportunity to ask questions, and then grade them when they’re done. It’s no wonder they fail. Most humans thrown into the deep end like that would fail too, which is exactly why a good manager would never do that to a new employee. New hires get onboarding, a colleague to ask, and feedback long before the deliverable is due.
Harvard Business Review published research in May arguing you shouldn’t treat agents like employees. The current test is something stranger: we give agents the employee’s assignment without any of the employee’s support, then read the result as a verdict on the technology.
The mistake isn’t optimism about the models. It’s treating the agent like a seasoned senior employee, the kind who can take a brief and disappear for a week, when today’s agents are closer to the greenest new hire you’ve ever worked with.
The test a junior hire could pass
Plenty of knowledge work is a briefed project. For a lot of people, that’s most of the job, and agents will eventually have to handle it. But nobody starts a new employee there, so the question isn’t whether today’s agents survive the senior’s test. It’s what a good manager gives someone in their first week.
Not the quarter-defining project. The arriving work: the forty overnight emails, the routine replies, the follow-ups that slip, the invoice approvals, the scheduling. Work where the context comes attached (the thread is right there, with the history and the actual ask), the steps are bounded, and someone reviews what goes out. We’ve written before about who starts the work; this is the slice of it a junior can own from day one.
That suggests a better benchmark for where agents are today: can one clear the mundane slice of your inbox and free you for the higher-value work that needs your judgment? Agents pass that test now, with current models. No capability breakthrough required.
The ladder doesn’t stop there, and it shouldn’t. We’ll eventually have agents that can take a brief and disappear for a week. But we shouldn’t skip steps. An employee earns that autonomy in two ways. They accumulate context, months of watching how the org decides, what the customers care about, which exceptions matter. And they ask for help when they’re unsure instead of guessing. Agents need the same two things: somewhere permanent for what they learn to accumulate, and a way to raise a hand mid-task. Give them both and the four-day brief stops being unthinkable. Skip them and no model release closes the gap.
What this predicts
If this reading is right, it makes a testable prediction: agents will succeed first not where they’re most autonomous but where the work is small, frequent, and arrives with context, and the “stall” will end without any dramatic model release. The improvement will come from accumulated company knowledge and better ways to ask for input, not raw IQ. Consolidate what the agent learns, and tomorrow’s agent gets further on the same broad task than today’s did, running on the same underlying model.
It also explains the deployment data instead of lamenting it. Those small bounded jobs enterprises actually kept agents on, the routing and triage and reminders, aren’t evidence of failure. They’re the junior work, done well: externally triggered, narrow, reviewable. The mistake is reading them as a consolation prize on the way to the briefed-project grail, when they’re the first rungs of the same ladder every employee climbs.
This is the same conclusion we keep arriving at from different directions. The second wave of software is fit-for-purpose AI apps rather than general chat, and fit-for-purpose, for an agent, starts with being stationed where the work shows up.
The bet we’ve made
this+that is built on this reading. The agent sits inside the message stream, across email and team chat, and the work triggers it. It reads what arrives, turns commitments and requests into tracked tasks, drafts what can be drafted, and runs workflows you describe in plain language. You don’t brief it, because the work already did. And the two things the ladder needs are the two things we’ve built around it. A workflow can pause mid-run and wait for a person before continuing, which is the early, limited form of an agent raising its hand. And the brain, our knowledge layer, is where what the agent and the team learn accumulates, which is how an agent gets more steeped in the org the longer it works there.
When agents do break through at scale, I don’t think it will be because the models finally got smart enough to survive a four-day brief delivered cold. It will be because we started them on junior work, where the work itself does the briefing, and let them earn their way up.
Key takeaways
- Zuckerberg’s leaked memo says agents haven’t progressed as fast as hoped, and the headline numbers (4% of multi-day tasks at human standard, near-universal deployment but mostly on small jobs) seem to back him up.
- Those numbers only spell “stall” if the benchmark is fair, and it isn’t: real freelance briefs delivered alone, with no one to ask and no company knowledge to draw on. Conditions only a well-onboarded senior employee could succeed under, applied to a day-one junior.
- The right test for today’s agents is the junior’s work: the mundane arriving stream, where context comes attached, steps are bounded, and a human reviews what goes out. Agents pass that test with current models, and clearing it frees people for the work that needs their judgment.
- Don’t skip steps. Agents earn autonomy the way employees do, by accumulating the organization’s knowledge and by asking for input when unsure. Consolidate what they learn, and tomorrow’s agent gets further on a broad task than today’s, even on the same model.
- this+that is built for the climb: the work triggers the agent, workflows can pause for a person mid-run (the early form of asking for help), and the brain accumulates what the agent and the team learn.