Leadership in the Age of AI

AI scores 90% on the benchmarks and finishes 2.5% of the real work. Measure the gap that matters.

Thomas Green 11 August 2026 6 min read
Key points
  • AI beats humans on coding tests, passes professional exams and tops reasoning benchmarks, which creates pressure to assume it can therefore do the actual jobs.
  • New evidence puts a number on the gap. The Remote Labor Index tested AI agents on 240 real, paid freelance projects; the best agent completed just 2.5% end-to-end, to a standard a client would accept.
  • The same models that score 80-90% on research benchmarks perform near the floor on real, economically valuable work. Benchmark power is not delivered value.
  • The reason: benchmarks test bounded questions with clean answers; real projects are ambiguous, need judgment, missing context and iteration to a payable standard. AI is strong at the first and weak at the second.
  • The move for leaders is to stop buying AI on the leaderboard and start measuring it on your own work: completion rate on real tasks, to your real standard, without a human quietly finishing the job.

You keep reading that AI now outperforms people on the hardest tests we have. It writes code that passes the benchmarks, clears professional exams, and tops the reasoning leaderboards. The natural inference, and the one your board may already be making, is that if it can ace the test it can do the job. Then you point it at your actual work and it stalls. The dazzling 80% demo does not become an 80% delivery. That gap is not your implementation falling short. It is the difference between scoring well on a test and finishing a real piece of work end to end, and there is now hard evidence for how wide it really is.

Researchers at Scale AI and the Center for AI Safety built the Remote Labor Index to measure exactly this. They took 240 real, paid freelance projects from Upwork, spanning software development, design, architecture and data, and tested whether AI agents could complete them end to end, to the standard a paying client would actually accept. The best-performing agent completed 2.5%. Not twenty-five percent, two-and-a-half. The same class of models that scores 80 to 90% on research benchmarks performs, in the authors' words, near the floor on real economically valuable work. The benchmark measures isolated skills in ideal conditions; the index measures delivered value in messy reality, and the two turn out to be worlds apart.

Why is there such a gap between the benchmark and the real job?

Because benchmarks test the part of work that is easy to test: bounded problems with a clean, checkable answer. Real projects are the opposite. The brief is ambiguous, the context is missing, the requirements shift, and getting it finished means exercising judgment, filling gaps, iterating with a client and pushing it over the line to a standard someone will pay for. AI is genuinely excellent at the bounded part and genuinely weak at the unbounded part. A high benchmark score tells you the model has raw capability. It tells you almost nothing about whether it can deliver your work, because your work is mostly the unbounded kind.

Measure AI on your work, not the leaderboard

The AI Strategy Session helps you find where AI genuinely finishes the job and where it only looks like it does, so you invest on evidence rather than hype, in ninety minutes.

Book your Strategy Session

What does this mean for how I invest in AI?

Stop buying on the leaderboard and start measuring on your own work. The demo that dazzles in a vendor pitch is the benchmark's cousin: a bounded, curated slice chosen to impress. What actually matters is the completion rate on your real tasks, to your real standard. That means running AI against a representative sample of the work you care about and measuring how much of it reaches done-and-acceptable without a person quietly finishing it, rather than how polished the first draft looks. The finish is where the value and the risk both live, and it is the thing a benchmark score will never show you.

So is AI overhyped?

No, and this is the balanced read that matters. The 2.5% is a snapshot of today, and the trajectory is upward: capability is climbing quickly and the completion rate will rise with it. The error is not believing AI is powerful, it plainly is. The error is assuming that benchmark power equals delivered value in your context right now. The sober position takes both seriously at once: treat the capability as real and rising, and treat the delivery gap as real and, for now, wide. Deploy AI fully where it genuinely finishes the job, and keep a human firmly in the loop everywhere the gap is still open, which today is most places.

  1. Measure AI on your work, not the leaderboard. Build a test set of real tasks and score end-to-end completion to your standard, not first-draft impressiveness.
  2. Separate "impressive draft" from "delivered result." The distance between them is where cost and risk hide. Track the finish, not the demo.
  3. Discount vendor benchmarks. A model topping a public leaderboard is a claim about capability, not a promise about your outcomes.
  4. Deploy where it finishes, augment where it stalls. Full automation only where completion is genuinely high; human-in-the-loop everywhere else.
  5. Re-measure on a schedule. The gap is closing, so today's "keep a human on it" becomes tomorrow's "let it run." Decide on evidence, not fear and not hype.
AI scores 80-90% on benchmarks and completes 2.5% of real freelance projects end to end. Benchmark power is not delivered value. Measure AI on your actual work, not the leaderboard.

What does this change for me as a leader?

It moves the question from "how capable is AI" to "how much of my actual work does it finish." The first has an impressive, rising answer. The second has a soberer one that should drive your decisions. The leaders who waste money are the ones who read the benchmark and restructure as if delivery had already arrived. The ones who get ahead measure the real completion rate on their own work, deploy hard where it is high, and keep human judgment in the loop where it is not, revisiting the line as the evidence moves.

This is the same discipline behind why most organisations fail at AI adoption: the constraint is rarely the model's raw power and almost always the gap between that power and delivered value in a real organisation. Close that gap with measurement and honest deployment, and you capture what AI can actually do today. Assume the benchmark is the business, and you will pay for a capability you have not yet got.

SourceFinding on AI capability versus delivered value
Remote Labor Index, Scale AI & Center for AI Safety (2025)Across 240 real, paid Upwork projects (software, design, architecture, data), the best AI agent completed just 2.5% end-to-end to an acceptable standard
Remote Labor Index (2025)Models scoring 80-90% on research benchmarks perform near the floor on real, economically valuable work
Remote Labor Index (2025)Research benchmarks test isolated skills in ideal conditions; real projects require end-to-end delivery to a payable standard, and the two diverge sharply
Implication for leadersMeasure AI on completion of your own representative work, not on public leaderboard scores or vendor demos

Frequently asked questions

Why does AI score highly on benchmarks but struggle with real work?
Because benchmarks test bounded problems with clean, checkable answers, while real projects are ambiguous, have missing context and shifting requirements, and need judgment and iteration to finish to a payable standard. AI is strong at the bounded part and weak at the unbounded part. The Remote Labor Index found the best agent completed just 2.5% of 240 real freelance projects end-to-end, even though similar models score 80-90% on research benchmarks. A high score reflects capability, not delivered value.
What is the Remote Labor Index?
It is a benchmark from Scale AI and the Center for AI Safety that measures whether AI agents can complete real, economically valuable remote work rather than isolated test questions. Researchers used 240 actual paid Upwork projects across software, design, architecture and data, and tested end-to-end completion to a client-acceptable standard. The highest-performing agent completed 2.5%, showing a large gap between benchmark performance and real-world delivery, though the figure is expected to rise over time.
How should leaders evaluate AI for their business?
Measure it on your own work, not on public leaderboards or vendor demos. Build a representative test set of real tasks and score how much AI completes end-to-end to your standard without a human quietly finishing it. Separate an impressive draft from a delivered result, deploy fully only where completion is genuinely high, keep humans in the loop where it stalls, and re-measure regularly because the gap is closing over time.
Thomas Green

About the author

Thomas Green

British technology futurist, AI keynote speaker and advisor. Thirty years across enterprise technology and AI strategy, helping leaders navigate the future of work. The futurist who died.

Get Thomas’ thinking into your inbox

Thoughts, stories and ideas on AI, leadership, and the future of work.