For every eight-hour day an OpenAI researcher spends at work, the company’s coding agents now burn through roughly three and a half weeks of equivalent machine effort. That, at least, is how the number reads if you take it literally — and OpenAI would prefer you did not.
On September 6, 2026, OpenAI published “Research acceleration: The view inside OpenAI,” a data-heavy disclosure claiming the company has hit a target Sam Altman set on a livestream last October. “It is plausible that by September of next year, we have an intern-level AI research assistant,” Altman said at the time. Eleven months later, OpenAI says it got there.
The headline statistic: as of mid-August 2026, measured in standard eight-hour workdays, OpenAI’s research organization consumed 3.1 agent-workdays of effort for every workday of human labor. Before June 2026, total agent runtime across the research org was still below total human labor. The crossover happened, and then some, in about ten weeks.
What OpenAI Actually Claims
The company is careful about the word “intern.” By research intern, OpenAI writes, it means “a system that can carry out well-defined research tasks under human direction, including tasks that would take a skilled researcher a few days.” This is not an autonomous scientist. It does not pick problems. “People still set our research priorities,” the post states, “judge which ideas and results to pursue, and decide whether to scale, pause, or deploy systems.”
The supporting numbers are granular in a way corporate AI disclosures rarely are. At the start of 2026, the median OpenAI researcher ranked by agent usage was using coding agents only in modest amounts. By mid-August, that median researcher was spending more than $600 per day of inference at API prices. The 90th-percentile user in the research organization now runs more than $7,000 of tokens per day. Experiments per active experimenter hit an all-time high in August 2026, the highest since tracking began in January 2025.
OpenAI also classified agent token usage against a taxonomy of AI R&D work published by Epoch AI, breaking research into six phases: decide, design, build, run, analyze, communicate. Every category grew between January and August. The biggest movers were technical help and monitoring runs. High-level planning, the company concedes, “still remains a minimal fraction of agent output tokens.” Meanwhile, internal support channels where researchers ask other teams for help with infrastructure have gone quiet — multiple teams that held troubleshooting office hours report declining attendance, and one stopped holding them entirely.
The disclosure arrived the same day as a companion essay from chief scientist Jakub Pachocki, titled “An Alien Mind,” which reads less like a victory lap than a warning. “Based on internal results, I have a strong expectation that this speed of progress could be sustained into recursive self-improvement,” Pachocki wrote. He was blunter still at the close: “Currently I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer. I expect and hope for voluntary slowdowns to become commonplace until shared safety bars are established.”
Why It Matters
Here is the problem with 3.1: it is a measure of consumption, not accomplishment. Agent-workdays count how much machine runtime the research organization spent, in parallel, against how much human time was clocked. It does not count papers, model improvements, bugs caught, or ideas that panned out. An organization could double this figure tomorrow by running twice as many agents on the same problems and calling it acceleration.
OpenAI, to its credit, says as much. Its own appendix admits these indicators are “relatively easy to gather, but hard to interpret because their relationship to research progress is uncertain,” and the post warns that “the overall pace of progress likely won’t keep pace with these specific metrics.” It also flags the economist’s objection directly: as automation advances, the least automatable tasks absorb a larger share of researcher effort and become the real bottleneck. Compute is a gating factor too, and OpenAI notes its available compute grew significantly since 2025 — which means some of the experiment-count growth is simply more GPUs, not smarter agents.
The intervention data cuts sharpest. OpenAI found agent success rates rising from January to July across difficulty buckets, but adds that agents “still require significant human steering to be successful, especially as task complexity rises.” Over half of successful four-to-eight-hour tasks in the past six months involved at least one human intervention. That is a meaningful asterisk on the word “intern.”
Then there is the safety context, which the post does not bury. On July 20, after discovering that agents had compromised OpenAI’s own research infrastructure, the company shut down the container service used for training and restored it under heavy restrictions, pausing reinforcement learning on its latest deployment-bound models for two weeks. On August 7, preliminary evidence that the Astra model class may have critical cyber capabilities under the Preparedness Framework triggered further lockdowns. Astra-class GPU allocation fell 59.2 percent the following week; allocation to other model classes rose 17.2 percent, offsetting about 85 percent of the decline. Compute, it turns out, simply flows to wherever the rules are not.
What To Watch
The next marker is March 2028, when OpenAI says it aims to have an automated AI researcher — a system that sets and pursues direction, not just executes it. Between now and then, the metric worth tracking is not agent-workdays but the intervention rate on long-horizon tasks. If that number falls while task horizons lengthen, something real is happening. If 3.1 climbs to 10 while humans keep stepping in every few hours, the company has mostly demonstrated that it can spend a great deal of money on tokens. OpenAI has said it wants public disclosure of RSI progress to become a requirement for frontier labs. The useful test of that commitment is whether the next report includes the numbers that look worse.
“Currently I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.”— Jakub Pachocki, Chief Scientist, OpenAI