← All Terms

Capability Threshold

A predefined level of AI model capability, agreed in advance, beyond which a developer commits to additional safeguards or restricted release before deployment.

Governance & Risk

A capability threshold sets the trigger, not the response. It is a specific point on a specific evaluation, for example a model’s performance on tasks related to biological weapon design, cyberattack automation, or autonomous self-replication, that a developer commits to treating as a decision point rather than a data point. Crossing it does not mean the model is banned. It means a pre-agreed process, additional red-teaming, restricted access, or a delay pending mitigation, is supposed to activate automatically instead of being decided case by case under commercial pressure.

The Seoul AI Safety Summit’s Frontier AI Safety Commitments made this the centrepiece of the voluntary regime: twenty companies agreed to publish their own thresholds and their own remediation plans for severe risk. The design has an obvious limitation built in. Each company sets its own threshold, tests its own model against it, and decides for itself whether the threshold has been crossed, with no external verification required by anything currently in force.

That gap is what separates a capability threshold from a safety standard. A standard is enforced by someone other than the party it constrains. As of 2026, capability thresholds for frontier AI are not; they are commitments a company can, in practice, quietly revise the week before a model release, and no signatory to the Seoul commitments has an obligation to disclose whether that has happened.