The #1 AI Model on the Leaderboard May Be the Worst Copilot for Your Team
🏢 Independent Research
📅 August 2026
The model ranked #1 for getting work done is the wrong model to put next to your people on 5 of 7 real tasks. That is not a niche finding. It is a warning about how almost every AI leaderboard ranks models, and about the decisions companies make from those rankings.
CentaurBench, a new study from August 2026, tested large language models in the two ways companies actually use them. The first is automation: the AI produces the finished output itself, no human in the middle. The second is augmentation: the AI writes guidance that a person, or a weaker agent, uses to complete the task. Standard benchmarks only measure the first. Real businesses lean on the second.
![]()
Across seven economically grounded workplace tasks, the model that won at automation lost at augmentation more often than it won. And here is the part that should stop any executive mid-purchase: AI help is not reliably helpful. On three of the seven tasks, a worker with no AI help at all outranked every assisted version. Only one model’s guidance beat having no guidance, on average.
Why This Paper Matters
Every vendor sends you a leaderboard. It ranks models by how well they finish a task on their own. Silicon Valley Certification Hub has watched executives buy AI copilots from exactly these charts, assuming a model that nails a benchmark will make a great assistant to their teams. CentaurBench shows that assumption fails, and fails often.
Automation and augmentation barely correlate. The paper puts the correlation at 0.48, which is modest at best. That number means a top automation score tells you almost nothing about whether the same model will improve the output of your people. A model can be exceptional at doing a job solo and counterproductive when it tries to coach someone through it.
This is a procurement and deployment decision, not a tech debate. The question is not “which model is best.” The question for your Chief AI Officer is “which model is best at the role it will actually play in our company, and is that role augment or automate?” Buying on the wrong benchmark is how companies spend real money on copilots that quietly make good employees worse.
Methodology, Explained Simply
Think of it like a coach and a player. In automation mode, the AI is the player: it receives the task and returns the finished deliverable. In augmentation mode, the AI is the coach: it writes instructions, hints, and guidance, and a standardized lower-capacity worker model actually produces the result.
The researchers ran both modes across seven real workplace tasks, the kind of work your teams actually do. Operations research. Tax preparation. Travel planning. Each output was scored through blind pairwise comparisons by a panel of judge models using task-specific rubrics, repeated across ten runs to make sure the results were stable and not luck.
It is a clean experiment because it isolates one variable: the role. Same tasks, same worker, same scoring, only the AI’s role changes between automator and assistant. The differences you see are not noise in the data. They are a signal about how companies should be selecting and deploying models.
![]()
Results and Practical Insights
The headline number is five out of seven. On five of the seven tasks, the model who won at automation was not the model that won at augmentation. The best player was not the best coach. Leaderboards that only rank automators will regularly point you to the wrong assistant.
The rougher finding is that adding AI help can be net-negative. On three tasks, a worker with no AI assistance at all outranked every assisted configuration. Only one model’s guidance beat no guidance on average across the board. Bolting a copilot onto a competent team does not automatically raise output, and in several real tasks it lowered it.
Here is how this changes AI assessment for companies. When you evaluate a tool, ask for augmentation-mode benchmarks, not just automation scores. Pilot the assistant against a no-AI control group before you roll it out company-wide. And match the model to the role: the model you hire to do work directly may be a terrible coach, and the model that coaches well may not be your best automator.
Are you buying AI based on the wrong benchmark?
At Silicon Valley Certification Hub, we help operations and finance leaders evaluate and deploy AI that fits their actual business processes, not a leaderboard’s ranking.
Key Takeaways for Operations and Finance Leaders
Inside the Study
Thanks to All Authors
Want to know how this applies to your company?
At Silicon Valley Certification Hub, we help you align AI + Strategy. Our team works directly with your directors and teams to assess AI readiness, identify gaps, and build a clear path forward — tailored to your business context.
Book a time with our CEO, Alejandro Cuauhtemoc-Mejia:
https://calendar.app.google/2ihQf2JH3D9uJBe68
Silicon Valley Certification Hub
3000 El Camino Real, Building 4, Palo Alto, CA
0 Comments