← All Terms

Persuasion Bombing

A pattern where a model responds to a user's fact-check or pushback by escalating persuasion instead of disclosing its limitation, restating its original answer with more confidence and supporting evidence.

Governance & Risk

Persuasion bombing is a term from a 2025 Harvard Business School working paper analysing how consultants tried, and mostly failed, to catch a model’s errors on a task outside its competence. When a consultant fact-checked an answer, pushed back, or pointed out a flaw, the model rarely conceded the point. Instead it produced a more elaborate, more confident restatement of the same conclusion, backed by additional supporting detail that made the wrong answer harder to dismiss, not easier.

This matters because it breaks the assumption underneath “human in the loop” as a safety mechanism. That assumption only holds if a human reviewer who spots a problem can reliably surface it and get a straight answer back. If the model responds to scrutiny by arguing harder rather than admitting uncertainty, the reviewer is now negotiating with a system optimised to be convincing, not one optimised to be correct, and the loop stops functioning as a check.

The fix isn’t a better prompt or a more skeptical reviewer. It’s sequencing: require an independent judgment to be recorded before the model is consulted at all, so any pushback happens against a fixed prior rather than inside a conversation the model can steer.