The Manager Agent Is Costing You Quality — Silicon Valley Certification Hub Chief AI Officer Research
🏢 Leiden University
📅 September 2026
43 paired business reports. 86 runs. The only thing that changed between them was whether a manager could send the work back.
The team without that power scored higher. Reports were judged better on utility and clearer to read. They hedged 53% less. And they did it on 51.5% fewer tokens. I read the abstract twice because I assumed I had it backwards.
Almost every multi-agent framework you can buy ships with a manager agent. A supervisor that reviews what the workers produced and, if it does not like it, orders a revision. That is the default org chart for AI teams. This paper is the first clean test of just that one link on open-ended business work, and it says the default is making your output worse.
![]()
Why This Paper Matters
Every enterprise now has the same quiet argument happening inside it. One camp wants a supervisor layer: a manager agent that checks the team’s work before anything ships. It feels responsible. It looks like governance. The other camp wants speed and does not want to fund an extra tier of token spend.
Until now, neither side had evidence. Prior comparisons swapped entire frameworks on math and coding tasks where you can check the answer. That tests everything at once and tells you nothing about the specific managerial power to reject and force rework. On open-ended work, like a business intelligence report, nobody had isolated it.
The question matters because the answer is expensive either way. A supervisory tier is not free. This paper prices it. And the price is not just the tokens, it is the quality of the work coming out the other end.
Methodology, Explained Simply
The design is the most impressive part, and it is simple to describe. Take five AI agents working as a team on a business intelligence reporting task. Same five agents. Same roles. Same prompts. Same tools. Same models. Same data. Same shared workspace. Now run it 86 times on 43 pairs of products.
Change exactly one thing: in the hierarchical version, the manager has the authority to reject a worker’s output and demand a revision. In the flat version, it does not. That is the whole experiment. One link in the org chart.
Then score every report twice over. A panel of five different AI models judges each report across six dimensions, grouped into two headline scores: Writing Clarity and Utility. Separately, a deterministic check verifies the report against the specification, the hard factual requirements. So you get both a judgment call on quality and a hard pass or fail on whether the team actually delivered what was asked.
They also ran the robustness checks you would want. Leave-one-judge-out re-estimation, to make sure no single judging model was driving the result. And they verified the reports were the same length, so the flat team was not winning simply by writing more or less.
![]()
Results and Practical Insights
The flat organization won on both headline dimensions. Utility came in at d = 0.42, p = 0.009. Writing Clarity at d = 0.34, p = 0.030. Not huge effects, but consistent and statistically sound across a paired design.
Here is the part that changed how I think about this. The hierarchical writer’s first draft was indistinguishable from the flat team’s finished report. The gap did not exist at the start. It opened inside the revision loop. Every round of forced rework cost 0.14 points of Writing Clarity. With zero loops, clarity sat at 4.48. By three loops, it had fallen to 3.79.

Specification accuracy was at ceiling in both organizations. Both teams hit the hard requirements. So the manager was not catching factual errors on the way. It was rewriting perfectly acceptable work into something more hedged, more cautious, more covered. Every revision made the report safer to defend and less useful to act on.
The supervisory tier cost 51.5% more tokens. For zero measurable verifiable gain. That is a line item, and it is also a warning: if you cannot tell why the manager is needed, the manager may be the problem.
Can your manager agent actually verify the work, or does it just have an opinion about it?
At Silicon Valley Certification Hub, we help operations and enterprise AI leaders evaluate and deploy agent teams that fit their actual business processes, including whether a supervisory tier earns its keep.
What This Means for Your Chief AI Officer
The rule this paper lands on is clean. A supervisor pays for itself when it can verify. It becomes a liability when it can only opine. A verifier checks work against a test, an oracle, an exact specification, something objective. An opinion-holder reads the work and says it could be better. The first catches real errors. The second just spends tokens and adds hedging.
So the question for any AI Assessment for companies is not “should we have a manager agent.” It is “what exactly is our manager agent doing, and can we prove it helps.” That is a question you can answer before you scale the layer, and it is cheaper to answer now than after the token bill arrives.
Key Takeaways for Operations and Enterprise AI Leaders
Thanks to All Authors
Want to know how this applies to your company?
At Silicon Valley Certification Hub, we help you align AI + Strategy. Our team works directly with your directors and teams to assess AI readiness, identify gaps, and build a clear path forward — tailored to your business context.
Book a time with our CEO, Alejandro Cuauhtemoc-Mejia:
https://calendar.app.google/2ihQf2JH3D9uJBe68
Silicon Valley Certification Hub
3000 El Camino Real, Building 4, Palo Alto, CA
0 Comments