FRONTIER LAB WATCHReviewed 2 September 2026
Attribution graphs still need a semantic contract
Mechanistic interpretability can expose computational structure, but the map is partial and the meaning assigned to it remains a scientific claim.
Mechanistic interpretabilitySemantic compilationCausal evidence
PRIMARY SOURCE · 27 March 2025
Tracing the thoughts of a large language model ↗
What the laboratory reports
- Anthropic uses attribution graphs to trace some internal pathways involved in model behavior.
- The research shows examples where a model's stated explanation does not match the computation suggested by the tracing method.
- Anthropic describes present methods as partial, potentially affected by tool artifacts and expensive to interpret.
The mathematical problem
An attribution graph is an estimated causal graph under an imperfect measurement operator. The hard questions are identifiability, omitted-variable bias, stability under interventions and whether a human label preserves the feature's operational meaning.
GERO's proposed response
- Version every graph together with the tracing method, thresholds and model checkpoint.
- Attach a semantic claim to each interpreted feature and require intervention-based evidence where feasible.
- Propagate partial coverage: an explained subgraph cannot certify computation outside the measured region.
A falsifiable experiment
- Repeat the trace across paraphrases and nearby prompts.
- Intervene on candidate features and test predicted downstream changes.
- Measure graph stability and the fraction of the behavioral claim supported by traced paths.
Read the primary source
This article is an original analytical summary, not a republication. Read Tracing the thoughts of a large language model for the laboratory's complete claims, methods and context.
