Agentic AI at Work: What Production Data Says About the Shift From Chat to Delegation — Silicon Valley Certification Hub Chief AI Officer Research
🏢 Stanford University · Carnegie Mellon University · Duke University · OpenAI
📅 June 2026
The median person in a legal role at OpenAI produced 13 times more output tokens in June 2026 than they did in November 2025. Researchers produced more than 50 times as many. That is not a software engineering story. That is a complete rebuild of how non-technical work gets done, and it happened in seven months.
This paper is unusual for one reason: it is not based on a survey. The authors worked with real production telemetry from Codex, OpenAI’s agentic coding tool, covering millions of requests. They compared three populations: individual personal accounts, organizational accounts, and OpenAI’s own employees. No self-reported “how much do you use AI” questions. Actual usage.
For any company trying to figure out where it stands on AI, this is the closest thing to a scoreboard that exists right now.
![]()
Why This Paper Matters
Every board I sit with asks the same question in different words. Is this AI thing actually changing how work gets done, or are we all just paying for licenses nobody opens? Until now the honest answer was: we do not really know, because almost all the evidence came from surveys and vendor case studies.
This paper changes that. The authors measured the gap between conversational AI and agentic AI directly. Conversational AI answers questions. Agentic AI takes actions, runs multi-step workflows, and works on your behalf while you do something else. That distinction matters more than any model benchmark, because the second one is what actually replaces labor hours.
And the growth is not subtle. Active Codex users grew more than fivefold in the first half of 2026. The fastest growth came from outside the original developer audience. So the adoption curve is bending exactly where most executives assumed it would stall.
Methodology, Explained Simply
Think of it as two doors into the same building. Door one is ChatGPT, where you ask a question and get an answer. Door two is Codex, where you hand over a task, walk away, and come back to finished work. The researchers counted who walks through which door, how often, and what they carry out with them.
They built a privacy-protecting pipeline, meaning no individual’s raw data or personal content was exposed to the researchers. They worked with aggregate patterns and model-estimated task complexity. That is why this paper can exist at all, and honestly it is the reason the data is credible.
They looked at three groups side by side. Personal accounts, which represent individuals experimenting on their own. Organizational accounts, which represent companies actually deploying this. And OpenAI’s own workforce, which represents the extreme case of a company that has fully committed.
They also measured sophistication, not just volume. How many people run several agents at once. How many use shared instruction sets for complex workflows, called skills. How complex the requests have become over time. That is the part most executives will find most useful, because volume without sophistication is just expensive novelty.
![]()
Results and Practical Insights
Here is the number that stopped me. Among organizational accounts, 63.3% of all output tokens now come from Codex, the agentic tool, not ChatGPT. Organizations are already routing the majority of their AI output through agents. Only 17.3% of those users touch Codex at all, which means a small group of power users is generating most of the value.
That lopsidedness is the whole story. Adoption is not evenly distributed inside companies. It clusters. A handful of people figure out the workflow, and they outproduce everyone else by an order of magnitude.

Request complexity tells a similar story. The share of individual Codex users submitting at least one task that an experienced human would need more than eight hours to finish rose nearly tenfold since January. People are not asking better questions. They are handing over entire projects.
Sophistication is climbing too. More than 10% of users run three or more agents concurrently at some point each week. About 26.6% use shared skills for complex workflows. Those are the signals that separate a company that bought AI from one that redesigned work around it.
If agentic AI already produces most of your organization’s AI output, who owns the workflow redesign?
At Silicon Valley Certification Hub, we help operations, legal, and finance leaders evaluate and deploy AI that fits their actual business processes, not the demo version.

What This Means for Your Chief AI Officer
Two lessons, and the second one is uncomfortable. First, the shift from chatting to delegating is already happening inside organizations, not just inside AI labs. Second, the value is concentrated in a small number of people who changed their workflows. Buying seats did not do that. Training and process design did.
This is where a real AI Assessment for companies beats a tooling decision. You need to know which functions have power users, which have none, and what the power users are actually doing differently. That is measurable. Most companies just arn’t measuring it.
Key Takeaways for Operations and Legal Leaders
Thanks to All Authors
Want to know how this applies to your company?
At Silicon Valley Certification Hub, we help you align AI + Strategy. Our team works directly with your directors and teams to assess AI readiness, identify gaps, and build a clear path forward — tailored to your business context.
Book a time with our CEO, Alejandro Cuauhtemoc-Mejia:
https://calendar.app.google/2ihQf2JH3D9uJBe68
Silicon Valley Certification Hub — 3000 El Camino Real, Building 4, Palo Alto, CA
0 Comments