Leading people in the age of agents: what AI speeds up and what it hides
keywords: people management, design leadership, applied ai, career
I've been writing a lot here about AI inside the design process. Today I want to pull on the part no agent takes off your hands: the people.

There's a calculation a lot of leaders are quietly running right now. If one person with AI delivers what two used to deliver, then half the team has become a cost.
It sounds clever, and worse, there's serious research handing it ammunition. Except that same research, read all the way through, charges a price nobody put on the spreadsheet.
Let's start with the study, because it's too good to skip. Between May and July 2024, researchers from Harvard, Wharton, ESSEC and Warwick ran a field experiment inside Procter & Gamble with 791 professionals, people with more than ten years at the company, working on real product development challenges.
The design was randomized: half got AI, half didn't; some worked alone, some in pairs, always matching someone from R&D with someone from the commercial side. The result that traveled around the world was this one: people working alone with AI reached the same solution quality as pairs working without AI.
Before you buy the whole package, there's something the paper itself declares: four P&G employees are co-authors, and the company funded the Harvard institute that coordinated the research.
That doesn't throw the result away. The experiment was pre-registered and randomized, more rigor than I see in most headlines about AI and productivity. I just think it's fair for you to know who was in the room when the number came out.
And there's the part that almost never travels with that result. Without AI, the R&D folks proposed technical solutions and the commercial folks proposed commercial ones, each pulling toward their own silo. With AI, both sides produced balanced proposals, regardless of their background.
Participants also reported more positive emotion and less negative emotion than those who worked without it. And when the researchers broke the innovation process into pieces, they saw exactly where the machine moved the needle: it improved idea generation, shifting the quality distribution upward, while human judgment kept its value in choosing which idea stays alive.
Generating got cheap. Choosing stayed expensive. And choosing is, in plain words, the job of whoever leads.
The "fewer people" promise trips at the door: AI relieved the part that could already be shared out and left standing exactly the part that concentrates on leadership.
Are you sure the team got faster?
That's the second question in the calculation, and it's far more slippery than it looks. In 2025, METR ran a controlled experiment with sixteen experienced open-source developers, each working in their own repositories. Before starting, they predicted AI would cut 24% off task time.
After finishing, they still believed they had been 20% faster. The measurement said the opposite: tasks done with AI took 19% longer.

Before anyone uses that to declare that AI slows people down, there's a caveat METR published themselves. In February 2026, they reported that the second round of the study produced no reliable signal: too many developers refused to take part in tasks without AI, and between 30% and 50% of those who did agree admitted to leaving some tasks out precisely because they didn't want to give up AI on them.
Their reading today is that there probably was a speed gain since then, without being able to pin down the size: the second round's estimates pointed toward faster, and none of them were statistically significant.
What survives from both studies has less to do with speed and more with the distance between what a person feels they produced and what they actually produced. That distance shows up in both directions, and it's the flimsiest instrument a leader could pick to make decisions about a team.
The 2025 DORA report, from Google Cloud, surveyed nearly five thousand technology professionals: 90% use AI at work and more than 80% believe it made them more productive.
In the same report, 30% say they have little or no trust in the code AI generates. Perceived productivity and trust in the output walking in different directions, inside the same sample.
So what exactly are you sizing a team against? "The team feels like it's flying" doesn't cut it.
AI amplifies what's already there
DORA's central conclusion is the sentence I'd nail to every executive wall: AI works as an amplifier. It multiplies the strength of teams that already had good process and exposes the mess of teams that didn't.
The report's numbers show both sides at once: AI adoption appears linked to more delivery throughput and better product performance, and also linked to less delivery stability. More going out, more breaking, when there's no net underneath.
Translated into people management: if your team lacks clarity on priorities today, with AI it will produce the wrong thing faster and in greater volume. If decisions bottleneck at you today, with AI they'll bottleneck at you with a longer queue behind them.
A new tool on top of a badly run organization doesn't fix the organization, it accelerates the defect. The report is blunt about it: the bigger return doesn't come from the tools, it comes from working on the system around them.
The step that's disappearing
Now the price that never makes it onto the spreadsheet, and the reason I wrote this post. The person who made me see this properly was Matt Beane, a researcher at UC Santa Barbara, who spent more than two years inside operating rooms observing robotic surgery. The comparison he draws hurts.
In open surgery, the resident went to the table because the procedure needed them: four hands and two people, almost the whole time. With the robot, the experienced surgeon handles it alone at the console. The resident became optional, and optional in practice means absent.
Beane reports they ended up with ten to twenty times less practice, and that plenty of them finished their residency licensed to operate with the robot without having trained enough. The ones who genuinely learned did it by bending the rules: simulators, video, hidden practice, away from their supervisor's eye. He called it shadow learning.
Notice there's no villain in this story. The robot is better than the old method. The surgeon did nothing wrong. And a whole cohort of residents still came out without learning.
Swap the robot for an agent and the surgeon for a senior designer, and it's the same scene. That tedious task the junior used to pick up, the competitor sweep, the first batch of wireframes, tidying the file, transcribing research, was exactly where they learned to see patterns.
Today that same work comes back finished from the agent, and it comes back better than it would from a junior six months in. What vanished along with the task is the place where that junior learned to get good.
And it's already showing up in the labor market. A study from the Stanford Digital Economy Lab, by Erik Brynjolfsson, Bharat Chandar and Ruyu Chen, found a 16% relative decline in employment for people aged 22 to 25 in the occupations most exposed to AI, such as software development and customer service, while employment for more experienced professionals in those same occupations held steady.
The decline concentrates where AI automates the task and all but disappears where it merely supports it. The authors themselves pump the brakes on the reading: with the broadest set of controls, the relationship between AI exposure and employment decline only becomes significant from 2024 on, so part of the earlier movement has other causes.
Cutting juniors is the most expensive saving there is, because who drops off the ledger today is the senior you won't have five years from now. And nobody becomes senior by watching an agent work.
So what's left for leadership
What's left is what was always hard, minus the excuse of a full execution queue. What's left is measuring outcomes instead of sensations: the question in a 1:1 today isn't how much you produced, it's what held up, what you decided to throw away and why. Delivery volume became the easiest metric to inflate and the easiest to fool whoever is watching.
What's left is fixing the house before buying the tool. Clear priorities, written criteria, someone who owns the decision. Skip that and go straight to the AI license and you've bought amplification of a problem you already had.
What's left is designing the learning step on purpose, because it won't come back on its own. That means deliberately reserving work AI could do faster for the people still learning, putting humans reviewing humans, and protecting pairing time even when the spreadsheet calls it inefficient.
It is inefficient. It's also the only known way to manufacture judgment, and judgment is the input your operation will burn through entirely over the next few years.
And what's left is distributing context. The knowledge that used to live in three people's heads and spread by osmosis in the hallway now has to be written down, because that's what someone will use to judge whatever the agent handed back. Context became team infrastructure.
One caveat so nobody leaves here distorted: none of this argues for slowing adoption. The gains are real and measured, including the emotional gain, which I found the loveliest finding in the P&G study.
What I'm saying is that AI moved the part of the work that could be divided, and not the part that requires someone deciding, growing people and carrying the consequences. That part is still done by humans, and it got bigger.
See you next time!
If you want to go to the source:
- Dell'Acqua, F. et al. (2026). The Cybernetic Teammate: A Field Experiment on Generative AI and Teamwork. Organization Science. The experiment with 791 Procter & Gamble professionals (read it on INFORMS, free working paper version at NBER). The sample appears as 776 in the working paper and 791 in the final revised version.
- METR (2025). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. The 19%-slower study and the inverted perception (metr.org).
- METR (2026). We are Changing our Developer Productivity Experiment Design. METR explaining why the second round produced no reliable signal (metr.org).
- DORA / Google Cloud (2025). State of AI-assisted Software Development. Nearly five thousand professionals, and the AI-as-amplifier thesis (dora.dev).
- Beane, M. (2019). Shadow Learning: Building Robotic Surgical Skill When Approved Means Fail. Administrative Science Quarterly (research summary from UC Santa Barbara).
- Brynjolfsson, E., Chandar, B. & Chen, R. (2025). Canaries in the Coal Mine? Six Facts about the Recent Employment Effects of Artificial Intelligence. Stanford Digital Economy Lab (study page, Nov/2025 PDF).
- Brynjolfsson, E., Chandar, B. & Chen, R. (2026). Canaries, Interest Rates, and Timing. The authors' own note where the relationship becomes significant from 2024 onward (Stanford Digital Economy Lab).
[ Read next ]

What is design? I came back to rewrite my own answer
I came back to rewrite my 2019 text. What changed? Empathy got more urgent in the age of AI, and data more important than ever.
[ 3 min read ]
From /imagine to a team of agents
From Midjourney to Claude: my journey with AI applied to product design, from asking a Discord bot for images to directing teams of agents today.
[ 5 min read ]
65% of Brazilians only access the internet through their phones. Is your website ready for the majority of the population?
Most Brazilians access the internet exclusively through their phones, and many pay for data by the megabyte. When we build websites fast with AI and only check them on desktop, that majority is exactly who gets left out. A conversation about responsiveness, page weight, and the real Brazilian internet.
[ 4 min read ]
[ loading comments... ]