Leading people in the age of agents: what AI speeds up and what it hides
keywords: people management, design leadership, applied ai, career
I've been writing a lot here about AI inside the design process. Today I want to pull on the part no agent takes off your hands: the people.

There's a calculation a lot of leaders are quietly running right now. If one person with AI delivers what two used to deliver, then half the team has become a cost.
It sounds clever, and worse, there's serious research handing it ammunition. Except the same experiment that hands over the ammunition also carries the number that knocks the math down, and that number almost never travels with it.
The study that arms the math, and the finding that disarms it
Between May and July 2024, researchers from Harvard, Wharton, ESSEC and Warwick ran a field experiment inside Procter & Gamble with 791 professionals working on real product development challenges. Experienced people: the study's summary table shows around ten to twelve years at the company on average, with a lot of spread — the standard deviation is over eight years.
The design was randomized: half got AI, half didn't; some worked alone, some in pairs, always matching someone from R&D with someone from the commercial side. The result that traveled around the world was this one: people working alone with AI reached the same solution quality as pairs working without AI.
That's the number feeding the math. Now the one that came out of the same experiment and stayed in the room.
The researchers isolated the solutions that landed in the top decile of the quality score. In the control group, people working alone without AI, 5.8% of solutions got there. Pairs with AI came in 9.2 percentage points above that, about 2.6 times the control’s chance (the paper rounds it to “around 3 times”), and that is the only one of the three effects that clears the statistical test: pairs without AI rose 3.7 points and people working alone with AI rose 1.9, and neither pulls away from chance.
And I'll take the honesty all the way, because this is where the math usually gets stretched. The comparison that holds is against the control. Between pairs with AI and pairs without AI, the difference gives p = 0.190. The data does not say AI is what carries the pair to the top; it says a pair with AI beats someone working alone without AI.

On the average, AI in an individual ties with the pair. At the top, it doesn't get there on its own. Cutting half the team buys the average and sells the top.
Fabrizio Dell'Acqua, first author of the study, put it this way in an interview with HBS Working Knowledge: if you want to empower an individual to be as effective as a team, give them AI; but if you want to be in that top 10% of performers, a full human team plus AI seems to be the recipe.
Who was in the room when the number came out
Before you buy the whole package, here's what the study itself declares. Four P&G employees are co-authors. The company made gifts to the HBS AI Institute, the Harvard institute that coordinated the research, between 2023 and 2025. And Karim Lakhani, one of the authors, was compensated as a consultant for P&G between 2021 and 2022. It's in the acknowledgements of the published version, alongside the sentence where they state they maintained full intellectual independence throughout the study.
The counterweight is public too, and it comes from inside: Ethan Mollick, a co-author, wrote that P&G had no control over the results or the data. That is the authors' word, not an independent audit.
There's one more thing the phrase "the design was randomized" hides. The workshop had 826 participants and the analyses use 791. The 35 left out weren't randomized: people who arrived late did the task alone without AI, and people whose seniority was above band 3 did it alone with AI. The analyses leave those people out, and the authors record that the results hold when they're included. It's in footnote 9.
None of that throws the result away. The experiment was pre-registered and randomized, more rigor than I see in most headlines about AI and productivity. The authors themselves list two limitations: it was a one-day virtual workshop, which doesn't capture the coordination and rework of a real team's routine, and the pairs always came from different functions, so groups with similar backgrounds or larger teams may behave differently.
Generating got cheap, choosing stayed expensive
And there's the part that almost never travels with the main result. Without AI, the R&D folks proposed technical solutions and the commercial folks proposed commercial ones, each pulling toward their own silo. With AI, both sides produced balanced proposals, regardless of their background. Participants also reported more positive emotion and less negative emotion than those who worked without it.
And there's the part holding up the rest of this post, so I'll give the address. The version published in Organization Science has a section the working paper didn't, 5.4.1, which breaks the innovation process into pieces. That was possible because the task had separate stages: each participant generated five ideas, picked one, and developed only the chosen one. Because the number of ideas was fixed at five for everyone, the researchers could separate AI's effect on idea quality from its effect on the choice.
The result has two halves, and only the first one travels. AI improved the average quality of the five ideas generated. When it came to picking which of the five to take forward, the people without AI got it right slightly more often: the authors record that human teams without AI showed a small advantage in selection accuracy, and sum up AI's role as a quality amplifier, and not as a decision enhancer.
The caveat is theirs and I want to repeat it: the selection advantage is small, the gain in generation more than compensates — the final outcome is still better with AI — and they themselves note that the people with AI may simply not have used it to choose. It isn't that the machine chooses badly. It's that nobody measured it choosing well.
Generating got cheap. Choosing stayed expensive. And choosing is, in plain words, the job of whoever leads.
The "fewer people" promise trips at the door: AI relieved the part that could already be shared out and left standing exactly the part that concentrates on leadership.
The step that's disappearing
Now the price that never makes it onto the spreadsheet, and the reason I wrote this post. The person who made me see this properly was Matt Beane, a researcher at UC Santa Barbara, who spent more than two years inside operating rooms observing robotic surgery. The comparison he draws hurts.
In open surgery, the resident went to the table because the procedure needed them: four hands and two people, almost the whole time. With the robot, the experienced surgeon handles it alone at the console. The resident became optional, and optional in practice means absent.
Beane reports they ended up with ten to twenty times less practice, and that plenty of them finished their residency licensed to operate with the robot without having trained enough. The ones who genuinely learned did it by bending the rules: simulators, video, hidden practice, away from their supervisor's eye. He called it shadow learning.
Notice there's no villain in this story. The robot is better than the old method. The surgeon did nothing wrong. And a whole cohort of residents still came out without learning.
Swap the robot for an agent and the surgeon for a senior designer, and it's the same scene. That tedious task the junior used to pick up, the competitor sweep, the first batch of wireframes, tidying the file, transcribing research, was exactly where they learned to see patterns.
Today that same work comes back finished from the agent, and it comes back better than it would from a junior six months in. What vanished along with the task is the place where that junior learned to get good.
And it's already showing up in the labor market. A study from the Stanford Digital Economy Lab, by Erik Brynjolfsson, Bharat Chandar and Ruyu Chen, found a 16% relative decline in employment for people aged 22 to 25 in the occupations most exposed to AI, such as software development and customer service, while employment for more experienced professionals in those same occupations held steady.
The decline concentrates where AI automates the task and all but disappears where it merely supports it. The authors themselves pump the brakes on the reading: with the broadest set of controls, the relationship between AI exposure and employment decline only becomes significant from 2024 on, so part of the earlier movement has other causes. Hold the two windows together: the 16% covers a stretch that starts before 2024, and the piece where the link to AI shows up clean is 2024 onward.
And there's a published challenge I have to bring in, because it hits what I'm arguing. The Economic Innovation Group, in Looking for the Ladder (Iscenko and Millet, January 2026), attributes much of the movement to the interest-rate cycle and shows that the drop in postings for the most exposed occupations began in March and April 2022, before ChatGPT. The part that costs me most is the second one: within those occupations, the EIG finds no evidence that junior postings fell more than senior ones. Jed Kolko, at Brookings, adds that the result shifts depending on which exposure measure you pick. I don't resolve this here. If the explanation turns out to be interest rates, my reading loses the best evidence it had, and what's left is what I see in practice.
Cutting juniors is the most expensive saving there is, because who drops off the ledger today is the senior you won't have five years from now. And nobody becomes senior by watching an agent work.
Are you sure the team got faster?
That leaves the third question in the math, and it's the slipperiest of the three. In 2025, METR ran a controlled experiment with sixteen experienced open-source developers, each working in their own repositories. Before starting, they predicted AI would cut 24% off task time.
After finishing, they still believed they had been 20% faster. The measurement said the opposite: tasks done with AI took 19% longer.

Before anyone uses that to declare that AI slows people down, there's a caveat METR published themselves. In February 2026, they reported that the second round of the study produced no reliable signal: too many developers refused to take part in tasks without AI, and between 30% and 50% of those who did agree admitted to leaving some tasks out precisely because they didn't want to give up AI on them.
Their reading today is that there probably was a speed gain since then, without being able to pin down the size: the second round's estimates pointed toward faster, and none of them were statistically significant.
What survives from both studies has less to do with speed and more with the distance between what a person feels they produced and what they actually produced. And that distance shows up in both directions. In the P&G experiment it showed up inverted: people using AI were 9.2 percentage points less likely to expect their own solution to land in the top 10%, while their objective score went up. They worked better and thought worse of themselves.
The 2025 DORA report, from Google Cloud, surveyed nearly five thousand technology professionals: 90% use AI at work and more than 80% believe it made them more productive. In the same report, 30% say they have little or no trust in the code AI generates. Perceived productivity and trust in the output walking in different directions, inside the same sample.
So what exactly are you sizing a team against? "The team feels like it's flying" doesn't cut it.
DORA's central conclusion is the sentence I'd nail to every executive wall: AI works as an amplifier. It multiplies the strength of teams that already had good process and exposes the mess of teams that didn't. The numbers show both sides at once: AI adoption appears linked to more delivery throughput and better product performance, and also linked to less delivery stability. These are correlations in a self-reported survey, and the delivery half flipped sign in a year: in the 2024 edition, adoption appeared linked to less throughput. The link to instability repeated in both. More going out, more breaking, when there's no net underneath.
Translated into people management: if your team lacks clarity on priorities today, with AI it will produce the wrong thing faster and in greater volume. If decisions bottleneck at you today, with AI they'll bottleneck at you with a longer queue behind them.
So what's left for leadership
What's left is what was always hard, minus the excuse of a full execution queue.
What's left is protecting the top, and not only the average. The top-decile finding is the most direct argument against dissolving a team into individuals with AI: if what you need is exceptional output, what the study measured as the recipe was a full human team plus AI. A whole team traded for individuals with an agent buys the average and loses exactly the band that justifies your function existing.
What's left is measuring outcomes instead of sensations. The question in a 1:1 today isn't how much you produced, it's what held up, what you decided to throw away and why. Delivery volume became the easiest metric to inflate and the easiest to fool whoever is watching. And the sensation errs in both directions: METR's developers thought they were 20% faster while working slower, and the P&G people thought worse of themselves while delivering better.
What's left is fixing the house before buying the tool. Clear priorities, written criteria, someone who owns the decision. Skip that and go straight to the AI license and you've bought amplification of a problem you already had. DORA is blunt about it: the bigger return doesn't come from the tools, it comes from working on the system around them.
What's left is designing the learning step on purpose, because it won't come back on its own. That means deliberately reserving work AI could do faster for the people still learning, putting humans reviewing humans, and protecting pairing time even when the spreadsheet calls it inefficient. It is inefficient. It's also the only known way to manufacture judgment.
And what's left is distributing context. The knowledge that used to live in three people's heads and spread by osmosis in the hallway now has to be written down, because that's what someone will use to judge whatever the agent handed back. If choosing is the expensive part, context is the raw material of whoever chooses. Context became team infrastructure.
The objection I can't resolve
If you're a director and you've read this far, you have a question ready for me, and it's a good one.
In surgery, residency exists to produce surgeons. It's regulated, it's the path required to register as a specialist, and somebody pays for it on purpose. Design has none of that. No company is obliged to train a junior, the benefit of training shows up five years out and probably at another company, because the person leaves before then. I'm asking your leadership to absorb a cost whose return the market captures.
That has a name: it's a collective action problem. Everyone gains if everyone trains, each one gains more if only the others train, and the result is that nobody trains. "Design the step on purpose" does not solve a collective action problem, and I'm not going to pretend it does.
What I have is smaller than the question, and I'd rather hand it over that way.
The return isn't all five years out. A junior trained in-house produces usable work long before becoming senior, and in the meantime they're the cheapest source of judgment your operation will have. The full return leaks; the partial return stays.
Whoever trained the person is also whoever has the best information about them and the first shot at keeping them. It's a weak argument, and it's still an argument.
And the honest one: what actually solves this is collective structure, the kind medicine has and design doesn't. A training pipeline funded by whoever captures the benefit, rather than the goodwill of one director at a time. Until that exists, training a junior is a bet made on a partial return, and I understand the people who don't take it.
Except the opposite decision has a price too, and it arrives later. If enough companies decide that training isn't their problem, five years from now everyone will be competing over the same shrunken stock of people with judgment, and the bill reappears on the salary line. For design that's my expectation, not a measurement. It isn't loose speculation, though: the abstract of Beane's article lists, among the outcomes of what happened in robotic surgery, hyperspecialization and a decreasing supply of experts relative to demand. The mechanism has been documented once already, in the one field where somebody went and looked.
One caveat so nobody leaves here distorted: none of this argues for slowing adoption. The gains are real and measured, including the emotional gain, which I found the loveliest finding in the P&G study.
What I'm saying is that AI moved the part of the work that could be divided, and not the part that requires someone deciding, growing people and carrying the consequences. That part is still done by humans, and it got bigger.
See you next time!
If you want to go to the source:
- Dell'Acqua, F., Ayoubi, C., Lifshitz, H., Sadun, R., Mollick, E., Mollick, L., Han, Y., Goldman, J., Nair, H., Taub, S. & Lakhani, K. (2026). The Cybernetic Teammate: A Field Experiment on Generative AI and Teamwork. Organization Science, published online 12 June 2026 (read it on INFORMS). This is the version carrying section 5.4.1, with the decomposition of the innovation process and the selection-accuracy measure (Figure 8), plus the acknowledgements with the funding and consulting disclosures. Open access.
- The free version is the NBER working paper, from April 2025, which circulates under a different title — The Cybernetic Teammate: A Field Experiment on Generative AI Reshaping Teamwork and Expertise — and with a different participant count: 811 at the workshop and 776 in the analyses, against 826 and 791 in the published version (NBER w33641). The figures this post uses — the years of tenure, the top-decile solutions, the inverted expectation — are identical across both; the decomposition of generating versus choosing exists only in the published one.
- Nover, S. (2025). When AI Joins the Team, Better Ideas Surface, HBS Working Knowledge, 17 October 2025 (read). The interview with Fabrizio Dell'Acqua, first author of the study, and the top 10% recipe.
- Mollick, E. (2025). The Cybernetic Teammate, One Useful Thing (read). Where one of the co-authors states that P&G had no control over the results or the data.
- METR (2025). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. The 19%-slower study and the inverted perception (metr.org).
- METR (2026). We are Changing our Developer Productivity Experiment Design. METR explaining why the second round produced no reliable signal (metr.org).
- DORA / Google Cloud (2025). State of AI-assisted Software Development. Nearly five thousand professionals, and the AI-as-amplifier thesis (dora.dev).
- Beane, M. (2019). Shadow Learning: Building Robotic Surgical Skill When Approved Means Fail. Administrative Science Quarterly (research summary from UC Santa Barbara).
- Brynjolfsson, E., Chandar, B. & Chen, R. (2025). Canaries in the Coal Mine? Six Facts about the Recent Employment Effects of Artificial Intelligence. Stanford Digital Economy Lab (study page, Nov/2025 PDF).
- Brynjolfsson, E., Chandar, B. & Chen, R. (2026). Canaries, Interest Rates, and Timing: More on the Recent Drivers of Employment Changes for Young Workers. Stanford Digital Economy Lab, 9 February 2026. The authors' own note where the relationship becomes significant from 2024 onward (Stanford Digital Economy Lab).
- Iscenko & Millet (2026). Looking for the Ladder. Economic Innovation Group, January 2026 (agglomerations.eig.org). The challenge attributing much of the movement to the interest-rate cycle. See also Kolko, J. (2026). Research on AI and the labor market is still in the first inning, Brookings, 10 March 2026 (brookings.edu).
Read next

What is design? I came back to rewrite my own answer
I came back to rewrite my 2019 text. What changed? Empathy got more urgent in the age of AI, and data more important than ever.
3 min read
From /imagine to a team of agents
From Midjourney to Claude: my journey with AI applied to product design, from asking a Discord bot for images to directing teams of agents today.
5 min read
Your site's first reader isn't a person
In a Pew study of 900 people, when Google showed an AI summary, clicks on a search result fell from 15% to 8%. Clicking a link inside the summary happened on 1% of visits. Whoever reads your site first is now a machine, it doesn't run your JavaScript, and what it understands is a design decision.
9 min read
loading comments...