Human-AI Interaction & Decision Quality
How effectively are humans and AI actually working together in your context? Benchmark your decision acceptance rates, automation bias exposure, and collaboration quality against global research from McKinsey, BCG, Stanford HAI, and MIT, adjusted for your industry, role, and decision type.
Listen to this briefing
Why the Centaur Model Beats Automation
Have your own measured numbers? Compare them against the benchmark ▸
Everything below defaults to published industry benchmarks for your selected industry, role, and decision type. If you have already instrumented your own acceptance and automation-bias rates, enter them here to see your actual numbers compared against the benchmark, not just the benchmark alone.
Decision Acceptance Funnel
Journey from AI recommendation → meaningful review → accepted → acted on.
Automation vs Augmentation Split
Optimal balance for this decision type. Centaur model (human+AI) outperforms either alone by –.
Decision Quality Dimensions
Five-dimension quality profile vs industry benchmark. Scores reflect improvement vs solo-human baseline.
Acceptance Rate by Industry (illustrative, directionally consistent with BCG / Stanford HAI research · optimal zone 65–85%)
This dashboard applies published third-party research (McKinsey, BCG, Stanford HAI, MIT, Edelman) to the industry, role, and decision type you select. It is not a measurement of your organisation's actual AI systems, and it is not legal, compliance, or governance advice. The recommended steps above are a starting point for your own instrumentation, not a substitute for measuring your own acceptance, bias, and quality metrics directly.
What each metric measures, what the research says, and how to improve your collaboration posture.
Decision Acceptance Rates & Automation Bias
Decision Acceptance Rate tracks the proportion of AI recommendations that humans act on. Automation Bias Index measures how many of those acceptances occurred without meaningful human review: the silent governance failure that most organisations have not yet instrumented. The EEOC and EU AI Act both require documented human oversight for high-risk decisions; automation bias is evidence that oversight is nominal rather than real.
- Sector pattern: acceptance and automation-bias rates vary sharply by sector. Healthcare and other consumer-facing use cases tend to show the highest acceptance of AI recommendations; legal and other high-liability domains tend to be the most conservative, since exposure to liability drives genuine review rather than rubber-stamping; financial services often carries the highest automation-bias risk, since credit and fraud decisions are scored at volume
- Optimal zone: a working range of 65–85% acceptance with under 25% automation bias is a reasonable target band. Below 50% signals undertrust in a system that has earned confidence; above 85% risks over-reliance
- Explainability effect: DARPA's Explainable AI (XAI) programme and the human-factors research that followed it consistently find that surfacing an explanation alongside a recommendation measurably reduces automation bias, though the size of the effect varies by study and domain
- Instrument your AI systems to log whether humans accessed the explanation before accepting. This is your automation bias rate
- Add mandatory explanation display before acceptance for P1/P0 decisions: interface friction that requires acknowledgement, not just click-through
- Set a review SLA: for high-stakes decisions, require logged time-on-task before acceptance (>60 seconds minimum)
- Report acceptance rates by team to leadership monthly. Regularly measuring and surfacing acceptance and bias rates is itself an intervention: teams that know they are being watched tend to self-correct faster than teams that are not
Automation vs Augmentation Spectrum
The automation vs augmentation split defines how AI is deployed across a decision portfolio. Automation means AI decides and acts without human involvement. Augmentation (the "centaur model") means AI advises, humans decide. The optimal split is not fixed. It varies critically by decision type, reversibility, regulatory context, and cognitive stakes. Getting this wrong in either direction destroys value: over-automation creates liability and error propagation; under-automation wastes the tool.
- Centaur model outperformance: the widely-cited BCG/Harvard/MIT/Wharton "Jagged Frontier" study of AI-augmented consultants found human+AI pairs completed 12.2% more tasks, 25.1% faster, with 40% higher-quality output than working solo, and that a clean human/AI division of labour ("centaurs") produced the highest accuracy of any configuration tested
- Directional pattern by decision type: the more routine and reversible a decision, the more automation a team can reasonably lean on; the more complex or consequential, the more the balance should shift toward augmentation, with high-stakes decisions kept almost entirely human-led and AI in an advisory role
- Governance framing: under NIST's AI Risk Management Framework, a fully-automated high-stakes decision with no meaningful human oversight is a governance gap worth raising with your risk committee, not a settled design choice
- Augmentation preference: industry surveys, including McKinsey's ongoing AI-adoption research, consistently find that most knowledge workers say they prefer AI as a thought-partner over a replacement decision-maker, with the preference strongest in professional and regulated fields such as healthcare and law
- Map every AI deployment to one of three tiers: Automate (routine, reversible, low-stakes), Augment (complex, consequential, regulated), Advise-only (irreversible, high-liability, safety-critical)
- Calculate your Return on Employee (RoE): measure hours freed from automated tasks + decision quality improvement per person. This is the centaur dividend
- Resist the automation bias in system design. The default should be augmentation, with automation requiring explicit justification and governance sign-off
- Survey team augmentation preference quarterly. Low preference scores predict adoption failure before it happens
Decision Quality & Cognitive Load
Decision quality in human-AI systems is multi-dimensional: accuracy (correctness), speed (time-to-decision), consistency (same decision in the same context), error rate (critical failures), and cognitive load (mental effort required). The composite score below is an illustrative index built from these five dimensions, not a certified external psychometric instrument: use it to track your own posture over time, not to benchmark against a named academic scale.
- Accuracy lift: most published human-AI collaboration research finds a meaningful accuracy improvement over solo human decisions, with the largest gains typically showing up in specialised, data-rich domains like healthcare and structured document review
- Error reduction: layering human review over AI output (or the reverse) consistently reduces critical error rates versus either working alone, though the size of the reduction is domain-specific
- Time-to-decision: routine decisions see the largest speed gains from AI assistance; high-stakes decisions see smaller gains, which is appropriate, since the goal there is better judgment, not just faster throughput
- Decision consistency: AI-assisted workflows tend to reduce the "decision fatigue" variance that affects human judgment later in a shift or a day, since a model does not tire the way a person does
- Trust in AI: Edelman's Trust Barometer research has tracked public trust in AI companies sitting roughly in the low-to-mid 50s out of 100 globally in recent years, with meaningful variation by sector and region
- Run a short quarterly survey across trust, transparency, control, and explainability. Even a simple 10-15 item internal instrument, tracked consistently over time, is more useful than a one-off benchmark against someone else's population
- Establish baseline accuracy and error rates before AI deployment. You cannot measure lift without a pre-AI baseline captured in the same period
- Track decision consistency using the same case presented to the same person twice over 4 weeks. The variance is your "human inconsistency baseline" that AI should reduce
- Monitor cognitive load as a leading indicator of adoption failure. High perceived effort predicts abandonment within 90 days, well before accuracy drops become visible

This one grew out of a pattern I kept seeing in the field: teams treating AI acceptance as a vanity metric, when the real signal is whether anyone actually reviewed the recommendation first. Adjust the industry and decision type and watch the automation bias number move independently of acceptance, that gap is where governance quietly breaks. The centaur model finding still surprises people every time: humans plus AI consistently beats either working alone. Good collaboration is measurable, not a feeling.
