← All Terms

Task Time Horizon

The length of a task, measured by how long a skilled human takes to complete it, that an AI agent can complete autonomously at a defined reliability threshold.

Agentic AI

Task time horizon reframes what a benchmark score is supposed to predict. Instead of asking whether a model gets a single question right, METR’s methodology asks how long a task can run, measured in the hours a skilled human would need to finish it, before the model’s success rate at completing it autonomously drops below a set threshold, typically 50%. A model with a one-hour horizon can be trusted to work unsupervised on something that takes a person an hour; hand it a task that takes a person a full day and its odds of finishing correctly collapse.

The reason this metric matters more than most public leaderboards is that it tracks over time on a strikingly consistent trend: METR’s March 2025 research found the horizon has roughly doubled every seven months for six straight years running, across very different model generations and labs. A single benchmark score tells you where a model stands today. A time-horizon trend tells you how fast the gap between “can answer a question” and “can run a multi-hour piece of work alone” is closing, which is the actual variable behind most claims about agents replacing sustained knowledge work rather than single tasks.

The metric also exposes a sharp reliability cliff most capability claims gloss over: at the time of METR’s measurement, frontier models were near-perfect on tasks taking a person under four minutes and under 10% reliable past four hours. That cliff, not the model’s best-case demo, is the number worth asking about before delegating any real task to an agent: not “can it do this,” but “at what task length does its reliability fall off.”