Lexicon · The risk of it

Mechanistic interpretability

Plain English. Reverse-engineering what is happening inside a neural network — identifying the internal features and circuits that produce behaviour — so a model's actions can be explained rather than merely observed. Advocates, including Anthropic's Dario Amodei, argue it is urgent and tractable; sceptics note that the celebrated results are on small models or narrow behaviours, and that nothing like a full account of a frontier model exists.

Why it moves money. Interpretability is the claimed answer to "can you trust the black box", which makes it the gate on regulated deployment — finance, medicine, defence — and a funded startup category, not just a research niche. If it scales, it becomes the audit layer every high-stakes deployment pays for. If it does not, black-box risk gets priced instead: heavier insurance, slower procurement, capped autonomy.

What to watch. The first case where an interpretability method predicts or prevents a production failure, disclosed in enough detail to check. That is the difference between an audit layer and a research aesthetic.

From the signals. Two ICML papers argue the visible reasoning trace is not evidence of alignment — the gap interpretability proposes to close from underneath.

Further reading. Dario Amodei, "The Urgency of Interpretability".

← All terms