Disagreement dossier 01Forecast F-01August 6, 2026

Capability / autonomy / the evidence gap

Can an AI do
a week of useful work?

The graph says: soon. The workplace asks a harder question—what counts as useful, reliable, and finished?

Test the claim →
01

The question beneath the question

Duration is not
the same as delegation.

An agent may persist for hours on a clean software task and still be unready for a week of work involving ambiguity, other people, changing goals, and costly mistakes.

That distinction matters to the singularity. The transition is not merely from answers to longer answers. It is from tools that respond to systems that carry consequential intent across time.

02

The measured horizon

How long can
an agent keep going?

METR defines a task horizon by how long the task takes a human expert—not by how long the model runs. Change the reliability standard to see why one headline number is not enough.

Required success rate
GPT-4oMay 2024
3 min
Claude Opus 4.5Nov 2025
1.1 h
Public frontierFeb–Mar 2026
1.5 h
F-01 thresholdby Dec 2028
40 h

Read carefully: the 2026 frontier values are approximate. Historical 80% figures are visual estimates from METR's published curves, included for context. F-01 is the Observatory's threshold, not a measured result. The horizontal scale is logarithmic.

03

The messiness test

Move from a test
into the world.

This explanatory model is not a reported benchmark. It reveals which assumptions disappear when “task completion” becomes “useful work.”

Illustrative dependable horizon8 h

Well-specified assignment

A human writes the brief and success criteria; tools and files are available.

  • Incomplete specification
  • Human dependencies
  • Changing requirements
  • Consequential errors
04

The disagreement

One trend.
Two readings.

The acceleration case

“Week-scale autonomy is the next visible step.”

  • Measured task horizons have lengthened rapidly across successive frontier models.
  • Tool use, coding, memory, and planning are being integrated into persistent agent workflows.
  • A specialized reimplementation benchmark has already produced estimates beyond 100 hours—evidence that long duration is technically reachable in bounded settings.
The friction case

“The benchmark is clean; work is not.”

  • At 80% reliability, the public frontier estimate is roughly 1.5 hours—not 12.
  • The main suite contains too few long, human-baselined tasks for confident extrapolation at week scale.
  • Changing requirements, coordination, judgment, and error recovery are the substance of many jobs, not incidental noise.
05

Observatory assessment

Likely in a bounded task.
Unproven in a living workplace.

The Observatory assigns a 62% probability that a frontier agent will reliably complete a bounded software or research task requiring a skilled human working week by the end of 2028.

This is not a forecast of one-week job replacement. It is a narrower, falsifiable threshold: an unfamiliar task, a declared success criterion, at least 80% reliability, and no human rescue.

62%Current · +8 points
from prior assessment
06

What would settle it

Evidence that
moves the number.

Raises forecast

Independent week-scale replication

At least 80% success on unfamiliar tasks across several domains, audited for hidden human help and benchmark leakage.

Lowers forecast

A persistent reliability ceiling

Long-horizon gains that disappear when specifications are incomplete, requirements change, or recovery from an early error is required.

Reframes forecast

A new definition of work

If useful output becomes a human–agent relay rather than autonomous completion, we will track that separately instead of moving the goalposts.

07

Evidence record

Read past
the headline.

These are the primary and independent reports used in this assessment. Measurements belong to their publishers; the synthesis and forecast belong to Mehan Observatory.

METR · May 2026Frontier Risk Report, February–March 2026

Public frontier: about 12 hours at 50% reliability and 1.5 hours at 80%; longer-task estimates remain uncertain.

METR · January 2026Measuring AI Ability to Complete Long Tasks: Time Horizon 1.1

The updated suite includes 228 tasks, but only five human-baselined tasks longer than eight hours.

METR · March 2026The Impact of Modeling Assumptions on Time Horizon Results

As a benchmark saturates, horizon estimates become more sensitive to statistical choices.

International AI Safety Report · 2026Extended Summary for Policymakers

Capability is improving, but remains jagged: systems can succeed at difficult tasks and fail at apparently simple ones.

Stanford HAI · 2026AI Index Report 2026

Coding performance rose sharply, while business deployment of agents remained in the single digits across nearly all functions.

METR · February 2026Update on Measuring AI Productivity

Real-world developer productivity remains difficult to estimate because adoption and selection effects complicate the comparison.