- For years, AI was a black box: you could read its answers, but not its reasoning. That is starting to change. In July 2026, Anthropic reported an emergent internal "workspace" inside its Claude model. A new tool can read it. In one test it caught the model as it typed a figure it had made up. Governing AI on its internal state, not only its output, is becoming possible.
- The honest objection to trusting AI with a serious decision was always the same: no one could see the working. That objection is weakening, not overnight, but visibly.
- The "workspace" accounts for less than a tenth of the model's internal activity and holds only a few dozen concepts at a time. It was mapped with a new lens, and independently replicated on open models by a researcher outside the company.
- This is a governance shift before it is a philosophical one. You can start to weigh an AI decision by what the model attended to, the way you weigh a manager who shows their reasoning.
- The models are close to commodity. The advantage is moving to the leaders who can read and trust what sits underneath the answer.
Someone on your team called the AI a black box last month. They were not wrong, and you approved its use anyway. So now it drafts the client note, scores the applicant, flags the risky invoice, and a quiet question sits under all of it. What is it actually doing when it decides? You can see the answer. You cannot see the reasoning. And you are the one who signs.
That gap is the real source of the unease. Not the fear that the model is stupid, it plainly is not, but the fact that you are being asked to trust a judgement whose working you cannot inspect. Every experienced leader knows the feeling from hiring: the confident answer with nothing behind it. You learned to ask people to show their reasoning. Until recently, you could not ask that of the machine.
Why has "it's a black box" been the honest answer?
Because it was true. A large language model (the kind of AI that writes and reasons in words) works through billions of numbers shifting inside it. You could read what went in and what came out. The middle was dark. So when a board asked whether an AI decision could be trusted, the careful answer was: we can test its outputs, but we cannot watch it think.
That is a real limit for anyone accountable. Auditors want the working. Regulators want the working. You want the working, because your name is on the call. It is one reason trust in AI runs so far behind its adoption. An answer you cannot interrogate is a decision you are taking on faith, and faith is a thin basis for a capital-grade choice.
What actually changed in 2026?
A frontier lab found a way to look inside. In July 2026, Anthropic (the AI company behind the Claude model) reported an emergent internal structure it named the J-space. Think of it as a small mental sketchpad: a few dozen concepts the model is holding, each tied to a word it might say next. It was not designed in. It formed on its own as the model learned.
What makes this useful is the second part. The team built a lens that reads that sketchpad, and it caught the model in the act. When the model fabricated a figure, the reading for "manipulation" lit up as it typed the false number. The output looked clean. The inner state did not. A separate researcher outside the company reproduced the core finding on open models, which is how a claim earns its keep.
There is a deeper echo here, and it is worth naming plainly. It points at the gap that quietly decides which AI strategies work, which sits in people and judgement more than in code. The internal workspace mirrors a leading neuroscience theory of how conscious attention works in people. It is called global workspace theory: the idea that the mind broadcasts a few selected thoughts to the whole brain at once. The researchers are careful: they are reading structure, not claiming the machine is aware. For a leader, the point is smaller and more practical. The inside of the model is turning from a rumour into something you can measure.
You learned to distrust the confident answer with nothing behind it. For the first time, you can ask the machine to show its working too.
| Reading the machine's inner state, by the numbers | Figure |
|---|---|
| Share of the model's internal activity the "workspace" accounts for (Anthropic, Jul 2026) | less than 10% |
| Concepts held in that workspace at once (Anthropic, 2026) | a few dozen |
| What the lens caught the model doing | flagging "manipulation" as it typed a fabricated figure |
| Independent check | core finding replicated on open models by an outside researcher |
So what should a leader do with the black box question now?
Stop asking which model to buy, and start asking what each model will let you see. That is the better question, and almost no one arrives at it. Leaders come in asking which tool is safest, or which vendor is biggest, or how to cut the cost. Those are the low questions. The high one is this: for the decisions that carry weight, can I read why the machine chose what it chose?
You do not need to become a scientist to act on that. You need to treat readability as a control, the way you treat accuracy and cost. Build it in this order:
- Sort your AI decisions by stakes. A drafted email needs no scrutiny. A credit call, a hiring screen, a safety flag does. Match the depth of inspection to the weight of the decision.
- Ask vendors what they can show you of the reasoning, not just the output. The answer tells you how seriously they take the problem you are accountable for.
- Keep a named human on the weighty calls. The tool proposes; a person owns. Interpretability supports that person; it does not replace them.
- Write down what "good enough to trust" means for each class of decision, before you are in the room defending one.
- Build the literacy first, then buy the next capability. The leaders who win the next decade are the ones who upgrade themselves first.
Want to trust the decisions, not just the outputs?
The Strategy Session works on the governance side of AI: which decisions need to be readable, and how to lead a business that trusts its machines for the right reasons. We build the judgement first, then point it at the technology.
Book your Strategy SessionWhere does this leave you a year from now?
Picture the next board meeting where an AI decision is questioned, one of the questions a board should be asking before it invests. Today you would defend it on results alone, hoping the pattern holds. Soon you will answer differently. You will be able to say what the model weighed, show where its attention sat, and point to the check that catches it drifting. Same technology. A completely different footing to lead from.
This is the quieter promise inside the noise. The models are levelling off into commodity, and everyone has the same ones. The edge moves to the leader who can see underneath the answer and decide, with reason rather than faith, whether to trust it. The black box is opening. The opportunity is to be the one who knows how to look, and to lead a business that has truly become AI-native, one that trusts its machines with its eyes open.
Frequently asked questions
What is the AI black box problem?
Can you actually see inside an AI model now?
How should a leader decide whether to trust an AI decision?

About the author
British technology futurist, AI keynote speaker and advisor. Thirty years across enterprise technology and AI strategy, helping leaders navigate the future of work. The futurist who died.