A control group is the half of an experiment that is left alone. One group of customers or cases gets the new model, nudge or process; a comparable group does not; and the difference in outcomes between the two is attributed to the change. Randomising who lands in which group is what makes the comparison fair, because it spreads everything else that could move the result (season, pricing, marketing, the economy) evenly across both.
It is the most defensible way to put a number on AI value, and the least used. Most AI business cases compare after with before, which credits the model with every other change that happened in the same period. DBS Bank measures the economic value it reports from AI largely this way; its Chief Data and Transformation Officer described the method as “a group of people gets AI treatment, a group doesn’t,” with the difference counted as value.
Where a control group is impossible, for example a fraud model that cannot ethically be withheld, the fallback is a documented baseline agreed before launch. Either way, the choice belongs in writing before the first result is reported.