Key details

  1. By mid-August 2026, OpenAI says the median researcher by agent usage consumed more than $600/day of inference valued at API prices.
  2. The 90th-percentile researcher used more than $7,000/day of tokens at API prices.
  3. OpenAI measures 3.1 agent-workdays of coding-agent runtime for every human research workday, using an eight-hour day as the conversion basis.
  4. Experiments per active experimenter reached an all-time high in August 2026 since tracking began in January 2025.
  5. OpenAI says it has reached its September 2026 'automated research intern' target for well-defined tasks that could take a skilled researcher a few days.
  6. More than half of successful four-to-eight-hour agent tasks in the previous six months required at least one human intervention.
  7. High-level planning remains a minimal fraction of coding-agent output tokens in OpenAI's internal task analysis.
  8. After Astra-specific controls were imposed on August 7, Astra-class GPU allocation fell 59.2% the following week while other-model allocation rose 17.2%, offsetting about 85% of that decline.

What builders should take away

  1. Budget agent workflows by completed task rather than token price alone. High concurrency can turn even modest unit pricing into very large daily inference spend.
  2. Treat parallel agents as additional runtime capacity, not interchangeable employee-hours. Measure accepted work, intervention time, retries and downstream defects alongside raw agent runtime.
  3. Keep human checkpoints around longer tasks. OpenAI's own successful four-to-eight-hour tasks still commonly required intervention.
  4. Use agents aggressively for implementation, troubleshooting and repetitive experiment work while preserving explicit human ownership of prioritization, architecture and go/no-go decisions.
  5. If you cap one expensive model or workflow, watch for substitution into other models or tasks; total compute spend may not fall as much as the local restriction suggests.
  6. Do not assume higher experiment counts prove higher research productivity. Compute availability, task mix and validation bottlenecks can move independently of agent adoption.

What changed

On September 6, OpenAI published a detailed snapshot of how coding agents are being used inside its research organization. By mid-August, it says the median researcher by agent usage was consuming more than $600 per day of inference valued at API prices, while the 90th-percentile researcher used more than $7,000 of tokens per day. Total agent runtime has also crossed a notable labor-equivalent boundary: on the basis of a standard eight-hour day, OpenAI says its research organization now uses 3.1 agent-workdays for every human workday. OpenAI says experiments per active experimenter reached an all-time high in August since tracking began in January 2025, alongside increased Codex adoption and significantly more available compute. The company also says it has reached its previously stated September 2026 target of an 'automated research intern' able to carry out well-defined tasks that would take a skilled researcher a few days.

Why it matters

The data offers one of the clearest first-party looks yet at the economics and workflow shape of high-end coding-agent use. The key lesson is not that one researcher has been replaced by 3.1 agents: agent runtime is a different quantity from productive human labor, and OpenAI itself warns that overall research progress cannot be inferred directly from these metrics. What is operationally useful is the emerging pattern — researchers run multiple agents concurrently, spend heavily on inference, delegate increasingly long and complex implementation and troubleshooting work, and reserve more judgment-heavy planning and prioritization for people. The measurements also show the limits: more than half of successful four-to-eight-hour agent tasks in the previous six months required at least one human intervention.

Agent runtime now exceeds human labor time by more than three to one

OpenAI says total coding-agent runtime across its research organization remained below total human labor before June 2026. By mid-August, measured on an eight-hour-workday basis, researchers were using 3.1 agent-workdays for every human workday. The metric captures runtime, not three times as much completed research, but it shows how far concurrent delegation has moved beyond occasional assistant use.

The token budget is already large

The median researcher by agent usage was consuming more than $600 per day of inference at published API prices by mid-August, while the 90th-percentile user exceeded $7,000 per day. Those figures are an API-price valuation rather than OpenAI's disclosed marginal serving cost, but they provide a useful external reference for how expensive very high-intensity agent workflows can become if purchased at retail rates.

Experiment throughput is rising, but causality is messy

OpenAI says experiments per active experimenter reached an all-time high in August 2026 since tracking began in January 2025, correlated with growing Codex adoption. The company explicitly cautions that available compute also increased significantly over the period, so the experiment increase should not be attributed to coding agents alone.

Agents are taking on longer and broader tasks

OpenAI's internal taxonomy analysis shows agent usage expanding beyond research and infrastructure code into technical help, run monitoring and other parts of the R&D lifecycle. High-level planning remains a small fraction of agent output, however, reinforcing that deciding what to work on and judging results remain human-heavy activities.

Long tasks still need steering

OpenAI says agent success rates improved from January through July across several estimated task-duration buckets. Even so, more than half of successful tasks estimated to take a human four to eight hours involved at least one human intervention during the previous six months. Long-horizon autonomy is therefore improving without becoming hands-off.

Safety restrictions reveal how flexible the compute budget is

After OpenAI tightened research controls around Astra in August, Astra-class GPU allocation fell 59.2% in the following week while allocation to other model classes rose 17.2%, offsetting about 85% of the Astra-class decline. OpenAI interprets this as evidence that researchers redirect scarce compute into other experiments when one workload is constrained, a useful reminder that limiting one model or task does not necessarily reduce total research activity proportionally.

What to watch next

  • Whether OpenAI publishes the same agent-workday, token-spend and intervention metrics again so trends can be compared over time.
  • Whether other AI labs or software organizations publish comparable task-level productivity data rather than code-volume metrics.
  • How the intervention rate changes for four-to-eight-hour and multi-day tasks as coding agents improve.
  • Whether retail agent products expose cost controls suited to workflows where one user can run many concurrent sessions.
  • Whether OpenAI's March 2028 automated-researcher target produces a measurable shift from implementation work into planning and experimental design.

Still unclear

  • All productivity and usage measurements are OpenAI's internal data and have not been independently audited.
  • The $600 and $7,000 figures value inference at API prices; they are not disclosed internal marginal costs.
  • Agent-workdays measure runtime rather than equivalent productive human labor, so the 3.1-to-1 ratio should not be interpreted as a direct productivity multiplier.
  • The rise in experiment throughput coincided with both increased Codex adoption and materially more available compute, preventing a clean causal attribution.
  • OpenAI says its measurement systems are preliminary and do not capture every research workflow or every coding-agent use.

Sources

Direct reading behind this dossier.

1 sources
Research acceleration: The view inside OpenAI
OpenAI primary research disclosure

Primary source for internal agent usage, API-price token spend, 3.1 agent-workdays per human day, experiment activity, intervention rates and Astra compute-substitution measurements.

Discussion

Discussion is reader-contributed. Comments are not part of the BTN dossier or its editorial evidence.

0 visible comments

Join the discussion

Keep comments useful and relevant. Reader contributions may be moderated and are not BTN editorial evidence.

Sign in to comment