← All Articles

AI Could Do the Job Tomorrow. The Math Still Caps How Much You Can Automate.

29 August 202613 min readAI StrategySharePDF

Listen to this article

AI Could Do the Job Tomorrow. The Math Still Caps How Much You Can Automate.

0:00

Executive Summary

Assume the capability question is settled: assume a model can already do most of the cognitive work currently done at entry, professional and even senior grades. That assumption still doesn’t tell you how much of your workforce you can actually replace. Five measurable constraints, oversight throughput, correlated model error, compute scarcity, legal attribution, and commoditised margins, set a ceiling on substitution that has nothing to do with what the model can do and everything to do with what your organisation can verify, staff, and afford to run. Most enterprise AI business cases size the automation. Almost none size the ceiling.

1/(f·v)

the corrected ceiling on automation speed-up, not the commonly cited 1/f. Adjudication effort (v) is a second, usually ignored lever

66%

of all AI compute will go to inference in 2026, up from 33% in 2023, Deloitte TMT Predictions 2026

2 votes

the effective independence of a 9-model LLM judge panel, per a 2026 correlated-error study, arXiv 2605.29800

7x

AI-exposed entry-level roles are now seven times more likely to require judgement and leadership skills once routine tasks are automated out, PwC 2026 Global AI Jobs Barometer

Core conclusions

  • The oversight bound most people cite, 1/f, is wrong. Adjudication is normally cheaper than production, so the real ceiling is 1/(f·v), and reducing v is an engineering problem most programmes never touch.
  • Redundancy from a second or third model buys far less assurance than headcount implies, because model errors are correlated, not independent.
  • Under equivalent AI capability available to every competitor, cost reduction transmits into price. The return sits with whoever holds a scarce complement, proprietary data, distribution, regulated position, not with substitution coverage itself.

The question everyone is answering wrong

Most debate about AI and jobs argues about capability: can a model do the work of a junior analyst, a paralegal, a claims adjuster. That’s the wrong question to build an operating model around, because even granting the most aggressive version of “yes” changes almost nothing about how much of that work you can actually hand off.

The useful question is different: if capability were fully solved tomorrow, what would still stop you from running your organisation at near-total automation? Run that constraint analysis and a specific, measurable ceiling appears, one set by oversight throughput, correlated error, the price of compute, legal attribution, and commoditised margins, not by what the model can technically produce. None of these constraints are exotic. All of them are already in your numbers, most companies just aren’t running them.

The formula everyone gets wrong

The standard claim is that if a fraction f of workload requires human adjudication, your maximum speed-up from automation is bounded by 1/f. At f = 0.10, that’s a ten-fold ceiling, and most planning stops there.

That bound is only correct if reviewing a decision costs the same effort as producing it. It almost never does. Adjudication is normally cheaper than production, and the corrected bound has a second term:

Maximum speed-up ≤ 1 / (f · v), where v is adjudication effort as a share of full production effort.

At f = 0.10 and v = 1.0, the ceiling is still tenfold. At f = 0.10 and v = 0.2, engineered so a reviewer isn’t redoing the work from scratch, it’s fiftyfold. That’s not a rounding difference. It’s the gap between an automation programme that plateaus and one that compounds, and it identifies two genuinely independent levers instead of one.

Reducing f is a model-reliability problem: better routing, narrower escalation classes, tighter confidence thresholds. Reducing v is a review-engineering problem: surfacing provenance and dissent signals up front, pre-computing the checks a reviewer would otherwise do by hand, ranking items by expected loss so attention lands where it changes the outcome. Almost every automation programme I’ve seen invests entirely in the first lever and never touches the second, which is the cheaper one to move.

The queue behind that ratio has its own failure mode worth naming: the adjudicator headcount needed is m ≥ (λ · f) / μ, sized to arrival rate, and it has to be sized against the ninety-fifth percentile of demand, not the average. A model defect typically raises both escalation volume and total inbound volume at once, which means sizing against the mean guarantees the first correlated incident produces a backlog exactly when you can least afford one.

Free tool

Board AI Oversight Checklist

Who actually checks the output, what happens when it’s wrong, and whether your board could explain the answer if asked. The v-lever above starts here, not in the model settings.

Redundancy doesn’t buy what you think it buys

The instinctive fix for an unreliable model is a second model to check the first one, or a panel of several. The assumption underneath that fix is that model errors are roughly independent, so agreement is meaningful and disagreement is a useful escalation trigger.

A 2026 study tested exactly this. Nine frontier judge models, drawn from seven different model families, were scored against 100 human annotations per item across three benchmarks. The panel’s nine votes behaved like roughly two independent votes: about three-quarters of the panel’s nominal independence disappeared because the models made the same mistakes on the same items. Panel accuracy fell 8 to 22 percentage points short of what genuinely independent voting would have achieved, and the single best judge in the panel matched or beat the full panel in every condition tested.

The mechanism is statistical, not incidental. For a population of n deciders with mean pairwise error correlation ρ, the variance of the aggregate error rate converges to ρσ² as n grows, not toward zero. Once the correlated component dominates, adding a fourth or fifth model buys almost nothing. Shared training data, shared architecture choices, and shared benchmarks push ρ up across model families in a way that shared procedure and shared incentives push it up across human reviewers in a single organisation, just not usually as high.

The fix isn’t more models. It’s architectural diversity: different input representations, different retrieval corpora, deterministic checks against invariants that don’t care which model produced the number, reconciliation against an independent data source. A ledger balance either reconciles or it doesn’t; that check doesn’t get weaker because the model that produced the entry was confident.

Compute is not getting cheap fast enough to matter

Business cases for automation nearly always assume the cost gap between machine and human labour is now large and only going to widen. That assumption quietly depends on inference staying cheap relative to demand, and the 2026 numbers argue the opposite.

Deloitte’s 2026 TMT predictions put inference, running the model, not training it, at roughly two-thirds of total AI compute this year, up from about half in 2025 and a third in 2023. Most of that inference still runs on expensive data centre silicon rather than cheap edge hardware, and capital is following: a new market for inference-optimised chips alone is projected past $50 billion.

That’s a Ricardian problem, not a technology problem. If serving capacity carries a positive shadow price because it’s scarce relative to demand, the decision an enterprise actually faces isn’t whether a model can do a task, it’s whether it should, given the opportunity cost of the compute that task consumes. Push automation demand up fast enough and the price of the constrained input rises with it, which erodes the very cost advantage that motivated the substitution in the first place. Comparative advantage doesn’t disappear when a machine gets absolutely better at everything; it reasserts itself through the price of whatever’s scarce, and right now that’s serving capacity, not human judgement.

Free tool

AI Trust, Risk & Governance Dashboard

Model the correlation and detection-rate assumptions above against your own bias, privacy and security thresholds, rather than trusting that a second vendor is buying you real redundancy.

Why maximum substitution isn’t maximum profit

Here’s the part most automation roadmaps skip entirely. Where equivalent AI capability is available to every competitor at comparable cost, the capability itself generates no rent. Cost reduction doesn’t stay in the P&L, it transmits into price, and the surplus lands with whoever holds a scarce complement to the capability, proprietary data with a legal right of use, distribution, a regulated position, physical assets, contracted serving capacity, institutional trust, not with the automation coverage number in the quarterly deck.

PwC’s 2026 Global AI Jobs Barometer, drawn from more than a billion job postings, shows this splitting the labour market in real time rather than compressing it uniformly. Roles PwC calls “professionalised”, where AI strips out the routine work and pushes judgement and expertise earlier into the career, are growing twice as fast and paying 42% faster wage growth than roles “democratised” by AI, where the tool just makes the role easier for a non-expert to do. AI-exposed entry-level roles are now seven times more likely to require the judgement and leadership skills that used to be reserved for senior grades, and postings for that seniorised type of entry-level role have grown 35% since 2019 while other entry-level postings fell 10%. Separately, Stanford’s Digital Economy Lab finds a persistent 13% relative employment decline for 22-to-25-year-olds in the most AI-exposed occupations, even after controlling for firm-level shocks, a gap that has continued to widen through mid-2026 even as overall employment shows no discernible AI effect.

Put those two findings together and the picture is not “AI replaces workers.” It’s “AI removes the apprenticeship rungs and pulls judgement earlier,” which raises the price of judgement precisely because it stops being trained as a by-product of routine production work. The rational objective for an enterprise isn’t substitution coverage. It’s return on whatever scarce complement it actually holds, with the automated capability treated as table stakes rather than advantage.

The measurement gap

Almost none of the constraints above show up in a standard automation dashboard, which tends to track coverage and cost per unit and stop there. A short set of leading indicators catches deterioration before it becomes a visible loss event, because the failure pattern in this domain is a slow drift in escalation rate, detection rate, and error correlation that precedes the incident by weeks, not a sudden capability failure.

Leading indicatorWhat it actually catches
Escalation fraction (f)Whether routing discipline is holding, or silently degrading
Adjudication effort ratio (v)Whether review is engineered to be cheap, or quietly re-doing the work
Detection rate on seeded defectsWhether your review process still catches what it’s supposed to catch
Error correlation across the model estateWhether a second model is real redundancy or the same mistake twice
Queue utilisation at the 95th percentileWhether the next correlated incident produces a backlog

Each of these needs a target, a threshold and a named owner before it’s worth anything on a dashboard. Report them to the same forum as the lagging cost and coverage numbers, not a separate one, because by the time the lagging numbers move, the leading ones already told you.

Where this leaves the roadmap

None of this is an argument against automating. It’s an argument against sizing an automation programme purely on model capability, because capability was never the binding constraint. Build the review engineering that shrinks v, treat a second model as a diversity decision rather than a redundancy checkbox, price the compute your automation actually consumes against the opportunity cost of the alternative, and go looking for the scarce complement your organisation actually holds before assuming the cost savings will stay yours. The enterprises that get this right in the next eighteen months won’t be the ones with the highest coverage number. They’ll be the ones who can tell you, with a straight face, exactly where their ceiling is and why.


Sources

  1. Brynjolfsson, Chandar, & Chen. (2025, updated 2026). Canaries in the Coal Mine? Six Facts About the Recent Employment Effects of Artificial Intelligence. Stanford Digital Economy Lab.
  2. Deloitte. (2026). TMT Predictions 2026: The AI Gap Narrows But Persists.
  3. Kohli et al. (2026). Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels. arXiv:2605.29800.
  4. Korinek & Lockwood. (2026). Public Finance in the Age of AI: A Primer. Brookings Institution.
  5. Information Technology and Innovation Foundation. (2026). AI Is Not Going to Reduce Labor’s Share of Income or Destroy the Tax Base.
  6. PwC. (2026). 2026 Global AI Jobs Barometer.
  7. Gallup. (2026). Organizational AI Adoption Jumps Six Points.

The AI Governance & ROI Executive Programme works through this constraint analysis line by line against your own escalation and error data, before the automation roadmap gets built, not after. Details are on the workshops page.

Was this useful?

Terence Kok
Before You Go

The number that changed how I model this wasn't a productivity stat. It was 2. Nine frontier judge models from seven different families, tested against 100 human annotations per item, behave like two independent evaluators, not nine, because they make the same mistakes on the same items. Every automation business case I've reviewed this year treats a second model, or a third vendor, as real redundancy. The data says it mostly isn't. If your assurance plan depends on models catching each other's errors, that plan is resting on a number close to one.

Terence Kok