What is Frontier Model?
Also called frontier AI.
A frontier model is one of the most capable general-purpose models available at a given time, typically trained at the largest scale currently feasible. The term is used in policy and safety discussion to designate systems whose capabilities are not yet well characterized and may pose novel risks. Its boundary is relative and shifts as the field advances.
Definitions vary and none is authoritative. Some proposals draw the line at a training compute threshold, which is measurable but becomes obsolete as efficiency improves. Others define it by demonstrated capability relative to existing systems, which tracks the concern better but is harder to apply consistently. Regulatory texts have used compute thresholds largely because they are administrable.
The category exists because of a specific worry: that the most capable systems may exhibit abilities their developers did not anticipate, including abilities with security or societal implications. This motivates pre-deployment evaluation, red-teaming, staged release, and incident reporting obligations that are considered unnecessary for smaller or narrower models.
The label is inherently transient. A model at the frontier when released is routinely matched by smaller and cheaper systems within a year or two, as training techniques, data curation, and distillation improve. Any regulatory or internal policy keyed to a fixed threshold therefore needs a revision mechanism or it will drift out of alignment with the actual risk landscape.
The framing is also contested. Critics argue that concentrating attention on the largest systems diverts scrutiny from the widely deployed smaller models causing measurable harm today, and that compute thresholds entrench incumbents by imposing compliance costs that only large developers can absorb. These are live policy disputes rather than settled questions.
Key points
- Designates the most capable general models of the moment
- Defined variously by training compute or demonstrated capability
- Motivates pre-deployment evaluation and staged release
- A moving target as efficiency improves each year
- The framing and its thresholds are actively contested
In practice
A governance policy requires a documented safety evaluation and a staged rollout for any model above a stated training compute threshold, and a lighter review below it. Two years later, a model well under that threshold matches the capability of one that originally exceeded it, because training methods improved. The policy's review clause triggers, and the threshold is revised downward rather than left to certify a capable system as low risk.