Which Work Tasks Should AI Do? A Cognitive Capability Framework
🏢 University of Cambridge
📅 August 2026
Every leader rolling out AI hits the same wall. Which tasks do we hand to the machine, which stay with people, and which get shared? Aggregate benchmark scores do not answer it. A model that crushes every public test can still stumble on your exact workflow, and by the time a vendor updates its numbers, the model has changed again.
This paper from a University of Cambridge team offers a cleaner way to think about it. Instead of asking “is this AI good?”, it asks “good at what?”. The authors profile AI systems and real jobs against the same set of core cognitive capabilities, then combine the two to score how suited a system is to a domain, a role, or a single duty. They validated it on six AI systems and 410 employees across six occupational domains.
![]()
Why This Paper Matters
Most AI assessment today is a guessing game. You read a headline score, run a small pilot, and hope it generalizes. The cost of being wrong is real. Tools get adopted that quietly fail on the tasks that matter, or good candidates get skipped because a benchmark did not reflect them.
This framework gives executives a repeatable answer. What surprised me most is that AI systems differ more across cognitive capabilities than across model families. Buying several models from the same vendor does not buy interchangeable skill. One model can be strong at language and memory but weak at reasoning about physical interactions, and that shape decides where it works.
For a Chief AI Officer, this is the difference between betting on hype and betting on evidence. It turns AI procurement into something you can measure and revisit as both models and roles change.
Methodology, Explained Simply
Think of it like a job description written in a language both humans and machines understand. The researchers defined eight core cognitive capabilities: language, memory, reasoning, social cognition, and a few more. This is the shared vocabulary.
On the AI side, they feed each model a benchmark battery where every question is tagged with the cognitive demands it makes. From how the model performs, they estimate its profile across the eight capabilities. On the work side, they ask domain experts to weight how important each capability is for their actual tasks. Both sides speak the same language, so the two can be compared directly.
Because the setup is modular, it stays current. A new model drops, you re-profile the system. A role changes, you re-weight the task. You never depend on a stale aggregate number.
![]()
Results and Practical Insights
The headline result is the one that should change how you select AI. Systems differed far more in capability across cognitive dimensions than across model families. Two models from the same family can have genuinely different strengths and weaknesses. Treating them as interchangeable is a mistake.
The workplace data converged too. The 410 employees clustered around a shared set of capabilities that drives most job fit. That is useful news: a relatively small cognitive core explains most of whether a task is automatable, so you do not need a complicated model of every job to get useful answers.
The team even ran the whole framework end to end on a real company, sorting tasks into a deployment grid. Tasks that are both important and well suited become prime pilot candidates. High importance but low suitability flags a capability gap, the honest signal that current systems are not ready and you should wait or steer elsewhere.
Are you choosing which tasks to automate based on a benchmark score, or on the actual cognitive demands of the work?
At Silicon Valley Certification Hub, we help operations and finance leaders evaluate and deploy AI that fits their actual business processes, not just the marketing claims.
What This Means for Your Chief AI Officer
This framework answers the question every Chief AI Officer keeps getting in the boardroom: which work do we actually hand to AI, and how do we know?
Silicon Valley Certification Hub has long argued that AI strategy fails less from lack of technology and more from weak decisions about where to apply it. This research is a practical tool for that gap. It gives you a scoping method to flag promising pilot candidates early and to steer away from tasks where current systems are unlikely to fit, before you spend budget on a pilot that cannot succeed.
It also supports the discipline of AI Assessment for companies: rather than a one time judgment, you build a living profile that updates as models improve and as roles shift. That turns AI adoption from a bet into a managed process.
Key Takeaways for Operations and Finance Leaders
Thanks to All Authors
Want to know how this applies to your company?
At Silicon Valley Certification Hub, we help you align AI + Strategy. Our team works directly with your directors and teams to assess AI readiness, identify gaps, and build a clear path forward — tailored to your business context.
Book a time with our CEO, Alejandro Cuauhtemoc-Mejia:
https://calendar.app.google/2ihQf2JH3D9uJBe68
Silicon Valley Certification Hub — 3000 El Camino Real, Building 4, Palo Alto, CA
0 Comments