Executive Summary
The advice given to knowledge workers worried about AI has mostly settled on one answer: don’t compete with the model, direct it. Learn to orchestrate agents instead of doing the task yourself. That advice has a shelf life shorter than the people giving it usually admit, because orchestration is itself a task, and the same capability curve that automated the task is already climbing toward automating the orchestration of it. Companies aren’t waiting to find out. They’re cutting the management layer first. None of this has a settled end state yet, which is exactly the argument for treating human inclusion as a design decision made now.
Every 7 months
the length of tasks frontier AI agents can complete autonomously has doubled on this cycle for six straight years, METR, March 2025
16% → 49%
automation potential of “managing and developing talent” work activities, 2017 estimate versus post-generative-AI 2023 estimate, McKinsey Global Institute
41%
of employees say their company has trimmed management layers in the past year, Korn Ferry survey of 15,000 professionals, via Fortune, June 2026
72%
of US managers believe their employees fear AI will make them less valuable at work, up 8 points year over year, Beautiful.ai 2026 manager survey
Core conclusions
- Companies are already removing the management layer that “orchestrate the agents” was supposed to become, not adding people into it. The org chart is answering the question before the reskilling advice does.
- The capability curve behind that decision isn’t a one-time jump. It’s compounding on a measured schedule, and orchestration sits squarely inside the category of task it’s climbing toward.
- Whether humans stay load-bearing in wisdom work is a design choice about where a system is built to require a person’s judgment, not an automatic consequence of how capable the models get.
Designing AI for human inclusion, ten slides
Save it, share it, or send it to whoever just got told to “learn to orchestrate the agents.”










The management layer being cut
The reskilling advice and the org chart are moving in opposite directions. Manager headcount at large public companies fell 6.1% between May 2022 and May 2025, with executive roles down 4.6% over the same stretch, according to Live Data Technologies data reported by Fortune. Estée Lauder and Match Group each cut roughly a fifth of their managers. Amazon eliminated around 14,000 corporate roles in October 2025, with CEO Andy Jassy explaining the logic directly: managers “want to put their fingerprint on everything,” he said, and the result is “people being in the pre-meeting, for the pre-meeting, for the pre-meeting, for the decision meeting.” Cloudflare cut about a fifth of its workforce in May 2026 while posting record revenue, and CEO Matthew Prince named exactly who the cuts targeted: “measurers,” his term for middle management, finance, legal, internal audit, and revenue recognition roles. A Korn Ferry survey of 15,000 professionals globally found 41% now say their own company has trimmed management layers in the past year.
None of these companies are framing the cuts as “we automated the manager’s job.” They’re framing it as removing a layer of coordination and approval that AI tooling made unnecessary. That’s a more precise and more uncomfortable claim than a simple automation story, because it means the job wasn’t replaced by a system doing the same tasks. It was removed because the reason for the layer to exist, keeping information moving and decisions checked between people who couldn’t otherwise see the whole picture, stopped requiring a person to do it.
The capability curve behind that decision keeps compounding
METR’s March 2025 research gives the clearest measurement of why this isn’t a one-off restructuring wave. Tracking task length by how long a skilled human takes to complete it, METR found the length of task frontier AI agents can complete autonomously at 50% reliability has doubled roughly every seven months, consistently, for six years. At the time of publication, models were near-perfect on tasks that take a person under four minutes, and under 10% reliable on tasks that take a person more than four hours. Orchestration, in the sense the reskilling advice means it, sits inside that longer-horizon band: deciding what to delegate, sequencing dependent steps, catching a bad output before it compounds. It’s exactly the kind of task the doubling curve is climbing toward, not a category exempt from it.
McKinsey’s own estimate of what generative AI changed captures the same shift in a single comparison. In McKinsey Global Institute’s pre-generative-AI 2017 estimate, “managing and developing talent” carried a 16% automation potential. By its 2023 reassessment, done after generative AI’s language and reasoning capabilities were factored in, that figure had risen to 49%. The activity most closely associated with what a manager does, not the spreadsheet work around it, is the one where the estimate moved the most.
Free tool
Human-AI Interaction & Decision Quality Dashboard
Benchmark decision acceptance rates and automation-bias exposure for your own industry and role against McKinsey, BCG, Stanford HAI, and MIT research.
Reskilling into orchestration assumes orchestration stays a human role
The advice to become the person who directs agents rather than the person who does the task isn’t wrong as a near-term move. It’s incomplete as a long-term one, because it treats “orchestrator” as a stable destination rather than another point on the same curve everything else has been sliding down. A market-entry analyst who spends two years learning to sequence agent workflows well is building a skill with the same shelf life as the task it was built to escape, if the pattern holds. The mid-level manager told to become a reviewer of AI output rather than a producer of it is being handed the review layer, which is precisely the layer Amazon and Cloudflare just cut because a person checking a person’s work, or a person checking an agent’s work, is the coordination cost both companies decided they no longer needed.
That’s the part the reskilling framing tends to skip. It answers “what should I learn” without answering “for how long does learning it matter,” and the honest answer right now is that nobody knows, including the people selling the framing with confidence. Beautiful.ai’s 2026 survey of 3,000 US managers found 72% believe their own employees fear AI will make them less valuable at work, up eight points from the year before, and 70% believe their employees fear AI will eventually cost them their job outright, up twelve points. Managers aren’t reporting confidence in the orchestration pivot. They’re reporting that the anxiety it was supposed to resolve is getting worse, not better, among the exact population being told to make that pivot.
Wisdom work was never only about task execution
The task-replacement framing that shapes most AI product design treats a manager’s job as a bundle of tasks: write the report, run the analysis, draft the review, schedule the sync. Automate enough of the bundle and the job looks solved. What that framing leaves out is the part of the job that was never a task in the first place: being the person whose judgment a decision is attached to, who can be asked why, who bears the consequence if the call turns out wrong, and whose presence in the room is itself what makes the outcome legitimate to the people affected by it. A capable agent can produce the analysis. It cannot, in any way that currently means anything to the people relying on the decision, be the party accountable for having made it.
That distinction is not sentimental, and it is not a permanent moat either. It is a design decision about where a system draws the line between output and accountability, and right now most AI systems are built to blur that line rather than mark it, because marking it is slower and less impressive in a demo. A system designed for task replacement optimises for getting a human out of the loop as fast as it can. A system designed for human inclusion has to decide, on purpose, which decisions keep a specific person load-bearing even when the system is capable of producing the same output alone, and build the workflow so that person’s judgment is structurally required.
| Design choice | Designed to replace the task | Designed to include the person |
|---|---|---|
| Default behavior | Produces the finished output, human review is optional | Stops at a checkpoint, human sign-off is structurally required |
| What gets optimised | Speed and completeness of the output | Speed of the output plus visibility of who is accountable for it |
| Where judgment sits | Assumed to be upstream, in training or prompt design | Kept downstream, at the point the decision lands |
| Failure mode | A confident wrong answer ships with no named owner | A slower answer, but someone specific can be asked why |
| What the org chart does over time | Removes the layer of people the system no longer needs | Keeps a smaller number of people whose sign-off the system cannot skip |

Coexistence needs specific design choices now
There is no clean end state to point to here, and it would be dishonest to pretend otherwise. Nobody currently building or deploying these systems has proof of where the automation frontier stops, or whether it stops at all before it reaches the accountability layer itself. What’s within reach is narrower and more useful: the systems being built and bought right now can be designed to keep a specific human decision load-bearing at a specific point, or they can be designed to route around that point the first time it’s inconvenient. That choice is being made today, by product teams and by the companies buying their tools, whether or not anyone frames it as a choice.
A small number of concrete practices follow from taking that seriously. Build the checkpoint into the workflow, not into a policy document, because a checkpoint someone can skip under deadline pressure isn’t a checkpoint. Name the person accountable for a given class of decision before the system ships. Measure how often the human checkpoint changes the system’s output, not just whether it exists: a rubber-stamp review that never overturns anything is a compliance artifact. Reject the framing that reskilling into “orchestrator” is a permanent answer, and treat every new capability jump as a fresh occasion to ask which decisions still deserve a structurally required human in them. None of this stops the capability curve. It decides what the curve is used for while it climbs.
The uncomfortable truth in the brief that prompted this post is the right one to sit with: we can train, upskill, and repurpose people faster than most organisations currently manage, and it may still not be enough, because the frontier keeps moving to wherever the retraining just arrived. That’s not a reason to stop trying. It’s the reason the harder, less comfortable work has to happen at the design table, deciding which decisions a system is built to need a person for.
Evidence & Methodology
Two of these come from direct measurement. One is an estimate McKinsey revised, not a count of what happened. The fourth is my own argument, not a data point at all.
| Claim | Source | Grade |
|---|---|---|
| The length of task frontier AI agents can complete autonomously has doubled roughly every seven months for six straight years | METR, March 2025 | Measured |
| Automation potential of “managing and developing talent” rose from 16% in 2017 to 49% in 2023 | McKinsey Global Institute, a modeled estimate, not an observed outcome | Estimate |
| 41% of employees say their company has trimmed management layers in the past year | Korn Ferry, 15,000 professionals surveyed, via Fortune | Measured |
| Whether humans stay load-bearing in this kind of work is a design choice, not an automatic result of how capable the models get | My own argument in this piece | My argument |
Sources
- METR. (2025, March 19). Measuring AI Ability to Complete Long Tasks.
- McKinsey Global Institute. (2023, June). The Economic Potential of Generative AI: The Next Productivity Frontier.
- Fortune. (2026, June 9). AI Agents Are Flattening Corporate Hierarchies, citing Live Data Technologies data and a Korn Ferry survey of 15,000 professionals.
- Beautiful.ai. (2026, April 22). AI’s Impact on the Workplace in 2026: 3rd Annual Survey of American Managers.
- Stanford Institute for Human-Centered AI. Human-centered AI, defining dimensions.
- Behaviour & Information Technology, Taylor & Francis. (2024). Understanding Human-Centred AI: A Review of Its Defining Elements and a Research Agenda.
If you’re deciding where a human decision needs to stay structurally required in an AI system you’re building or buying, that’s exactly the kind of design call my consulting work helps sequence.
