Capability / autonomy / the evidence gap
Can an AI do
a week of useful work?
The graph says: soon. The workplace asks a harder question—what counts as useful, reliable, and finished?
Test the claim →The question beneath the question
Duration is not
the same as delegation.
An agent may persist for hours on a clean software task and still be unready for a week of work involving ambiguity, other people, changing goals, and costly mistakes.
That distinction matters to the singularity. The transition is not merely from answers to longer answers. It is from tools that respond to systems that carry consequential intent across time.
The measured horizon
How long can
an agent keep going?
METR defines a task horizon by how long the task takes a human expert—not by how long the model runs. Change the reliability standard to see why one headline number is not enough.
Read carefully: the 2026 frontier values are approximate. Historical 80% figures are visual estimates from METR's published curves, included for context. F-01 is the Observatory's threshold, not a measured result. The horizontal scale is logarithmic.
The messiness test
Move from a test
into the world.
This explanatory model is not a reported benchmark. It reveals which assumptions disappear when “task completion” becomes “useful work.”
Well-specified assignment
A human writes the brief and success criteria; tools and files are available.
- Incomplete specification
- Human dependencies
- Changing requirements
- Consequential errors
The disagreement
One trend.
Two readings.
“Week-scale autonomy is the next visible step.”
- Measured task horizons have lengthened rapidly across successive frontier models.
- Tool use, coding, memory, and planning are being integrated into persistent agent workflows.
- A specialized reimplementation benchmark has already produced estimates beyond 100 hours—evidence that long duration is technically reachable in bounded settings.
“The benchmark is clean; work is not.”
- At 80% reliability, the public frontier estimate is roughly 1.5 hours—not 12.
- The main suite contains too few long, human-baselined tasks for confident extrapolation at week scale.
- Changing requirements, coordination, judgment, and error recovery are the substance of many jobs, not incidental noise.
Observatory assessment
Likely in a bounded task.
Unproven in a living workplace.
The Observatory assigns a 62% probability that a frontier agent will reliably complete a bounded software or research task requiring a skilled human working week by the end of 2028.
This is not a forecast of one-week job replacement. It is a narrower, falsifiable threshold: an unfamiliar task, a declared success criterion, at least 80% reliability, and no human rescue.
from prior assessment
What would settle it
Evidence that
moves the number.
Independent week-scale replication
At least 80% success on unfamiliar tasks across several domains, audited for hidden human help and benchmark leakage.
A persistent reliability ceiling
Long-horizon gains that disappear when specifications are incomplete, requirements change, or recovery from an early error is required.
A new definition of work
If useful output becomes a human–agent relay rather than autonomous completion, we will track that separately instead of moving the goalposts.
Evidence record
Read past
the headline.
These are the primary and independent reports used in this assessment. Measurements belong to their publishers; the synthesis and forecast belong to Mehan Observatory.
METR · May 2026Frontier Risk Report, February–March 2026Public frontier: about 12 hours at 50% reliability and 1.5 hours at 80%; longer-task estimates remain uncertain.
METR · January 2026Measuring AI Ability to Complete Long Tasks: Time Horizon 1.1The updated suite includes 228 tasks, but only five human-baselined tasks longer than eight hours.
METR · March 2026The Impact of Modeling Assumptions on Time Horizon ResultsAs a benchmark saturates, horizon estimates become more sensitive to statistical choices.
International AI Safety Report · 2026Extended Summary for PolicymakersCapability is improving, but remains jagged: systems can succeed at difficult tasks and fail at apparently simple ones.
Stanford HAI · 2026AI Index Report 2026Coding performance rose sharply, while business deployment of agents remained in the single digits across nearly all functions.
METR · February 2026Update on Measuring AI ProductivityReal-world developer productivity remains difficult to estimate because adoption and selection effects complicate the comparison.
The dossier changes when the evidence does
Follow the
disagreement.
Receive new dossiers, forecast revisions, and the evidence that changed our mind.