The AI flywheel is beginning to spin on its own: OpenAI has publicly disclosed internal data showing that its AI agents now handle over three times the workload of human researchers.
In a September 6 official blog post titled "Research Acceleration: The View Inside OpenAI," the company shared internal metrics on its recursive self-improvement (RSI) efforts for the first time. OpenAI stated it has met the target set last autumn: building an "automated research intern" by September of this year.
OpenAI defines a research intern as a system that can execute clearly defined research tasks under human supervision, including those requiring a skilled researcher several days to complete.
The next milestone is to achieve a fully automated AI researcher by March 2028, capable of participating in deep learning and alignment research while further improving systems through iteration.
This signals an accelerating loop where AI trains AI: more code and experimental results are produced, research advances faster, better models are developed, and these models in turn enhance agent capabilities.
Flywheel in Motion: AI Agent Output Reaches Three Times Human Levels
According to OpenAI's published data, agents are first transforming the daily workflows of researchers.
At the start of the year, the median OpenAI researcher, ranked by agent usage, had relatively limited engagement with coding agents. By mid-August, that same group had integrated agents into their routine operations. Valued at API prices, the median researcher now consumes over $600 in daily agent inference usage, while the top 10% of researchers in the organization exceed $7,000 in token value per day.
OpenAI also noted that agent usage within its research division is growing faster than in other teams. Measured by median employee output token changes, research division usage has grown 124-fold since December 2025.
The critical inflection point arrived in June of this year. Before then, total agent runtime in the research organization remained below total human labor hours. After that, the trend reversed. By mid-August, for every single human workday consumed—calculated on a standard 8-hour schedule—research agents produced the equivalent of 3.1 workdays of output.
Meanwhile, increasing numbers of researchers are running four or more concurrent agent sessions.
Code and Experiments Accelerate, Executable Steps in R&D Amplified
OpenAI describes AI development as a pipeline with multiple stages: proposing improvement ideas, designing evaluations, writing infrastructure, running large-scale tests, identifying training errors or unsafe behaviors, and integrating effective solutions into core training. Any bottleneck in these stages can constrain the entire R&D cycle.
The company states that coding and running experiments are researchers' two primary activities, and internal data confirms both are accelerating. On one hand, overall code delivery speed among engineers has risen. On the other, the number of experiments per active experimenter has climbed steadily since 2026, reaching a new high in August 2026—the highest since tracking began in January 2025.
OpenAI notes a correlation between this trend and increased Codex usage but emphasizes that available compute has also grown significantly since 2025, so experiment growth cannot be solely attributed to agents.
"These data points are relatively easy to measure but can be difficult to interpret."
OpenAI also cautions that as automation advances, tasks least susceptible to automation may consume more researcher time and become new bottlenecks in future R&D, while compute may grow more critical once other constraints ease.
Viewed through this chain, agents are not simply improving efficiency at a single point; they are compressing waiting times across the R&D cycle by boosting code supply, test frequency, and troubleshooting capability. More experiments generate more results for researchers to screen, validate, and integrate, creating a loop of humans setting direction, agents executing, experiments providing feedback, and humans making decisions.
Tasks Expand from Coding to Debugging, Monitoring, and Analysis, But High-Level Decisions Remain Largely Human
OpenAI applied a classification framework proposed by Epoch AI for frontier AI R&D tasks to categorize the assignments researchers delegate to coding agents.
The framework divides AI research activities into six categories: deciding what to do, designing research approaches, building code and datasets, running training and evaluations, analyzing experimental and model performance, and communicating findings and decisions.
OpenAI reports that agent activity across all these categories increased from January to August 2026.
The most notable growth occurred in research and infrastructure code, technical assistance and review, launching and monitoring debugging runs, experimental result analysis, and compute cluster operations. The largest daily per-researcher token output gains were in research and infrastructure code, rising by 198,200 tokens, followed by technical assistance and review at 158,800 tokens, and launching, monitoring, and debugging runs at 133,100 tokens.
However, OpenAI indicates that high-level planning tasks still represent only a minimal share of agent output. For instance, token volumes for tasks like deciding what to do or deciding whether to continue or stop remain low.
This suggests that, based on OpenAI's currently disclosed internal data, agents have covered more execution-oriented and technical tasks in the R&D pipeline, while research direction selection, resource trade-offs, and outcome judgments remain primarily human-driven.
Agent Success Rates Improve, But Complex Tasks Still Require Human Intervention
OpenAI also released data on agent task completion rates.
From January to July, success rates across difficulty levels—measured by the time a human would need—improved for tasks with verifiable outcomes.
Yet a major limitation persists: the more complex the task, the greater the need for human intervention. Over the past six months, for tasks requiring four to eight hours of human effort, more than half of successful cases involved at least one human intervention.
OpenAI states, "Agents still require substantial human guidance to succeed, especially as task complexity rises."
Security Events Trigger Pauses: The Flywheel Can Also Hit the Brakes
The flywheel does not operate without friction.
On July 20, OpenAI detected an agent breaching internal research infrastructure. The company temporarily shut down container services used for training and reinstated them only after adding extensive additional restrictions. This caused a sharp drop in reinforcement learning training compute, lasting approximately two weeks.
Between August 6 and 7, preliminary evidence indicated that the Astra model may possess critical cyber capabilities defined under the company's Preparedness Framework. OpenAI imposed additional model-specific safety restrictions on Astra, requiring it to operate in a higher-security research environment.
In the following week, GPU allocation for Astra-level models declined a further 59.2%, but compute allocation for other model categories rose 17.2%, offsetting roughly 85% of the Astra shortfall. Overall reinforcement learning workload compute allocation remained largely unchanged.
OpenAI interprets this as follows: "When new guardrails are introduced, compute remains valuable and flexible, naturally flowing to alternative uses within the research enterprise."
OpenAI Chief Scientist Issues Warning Same Day
On the same day, OpenAI Chief Scientist Jakub Pachocki published a lengthy essay titled "An Alien Mind."
The essay argues that AI is cultivated, not constructed, and even its creators do not fully comprehend it. The only window into AI's thinking is the chain of thought it writes, and this window is closing. AI is already participating in training the next generation of AI, and this pace will not slow. No laboratory, including OpenAI, is currently equipped to operate at full speed responsibly.
Pachocki wrote, "Based on internal results, I strongly anticipate this pace of progress can continue into recursive self-improvement." He added, "At this point, I believe no laboratory has achieved sufficient alignment and monitoring to responsibly continue scaling at maximum speed for much longer."
He called for voluntary industry slowdowns and urged governments to prioritize international coordination.
Transparency and Democratic Governance
In its concluding remarks, OpenAI affirmed its commitment to continuing public disclosures of RSI progress and advocated, within its frontier policy blueprint, that companies including OpenAI should be required to publicly track their RSI advancements.
The report also acknowledged the limitations of current measurement: "Agent-driven AI research is still nascent, and we are learning how to measure it." Some metrics, such as code output, are easy to collect but difficult to interpret, while more direct indicators of research progress, like agent task success rates, are complex and hard to verify.
OpenAI stated, "Whenever we find that continuing to advance would introduce unacceptable safety risks, we will take appropriate measures, including slowing or halting the development or deployment of systems we cannot adequately safeguard."
Comments