← Research index
FRONTIER LAB WATCHReviewed 2 September 2026
Attribution graphs still need a semantic contract
Mechanistic interpretability can expose computational structure, but the map is partial and the meaning assigned to it remains a scientific claim.
Xamit KadirbekovIndependent analysis · Source: Anthropic
Mechanistic interpretabilitySemantic compilationCausal evidence
STATUS · SOURCE REPORT + GERO ANALYSISThis brief has not independently reproduced the laboratory's experiment.
PRIMARY SOURCE · 27 March 2025
Tracing the thoughts of a large language model ↗
What the laboratory reports
Anthropic uses attribution graphs to trace some internal pathways involved in model behavior.
The research shows examples where a model's stated explanation does not match the computation suggested by the tracing method.
Anthropic describes present methods as partial, potentially affected by tool artifacts and expensive to interpret.
The mathematical problem
An attribution graph is an estimated causal graph under an imperfect measurement operator. The hard questions are identifiability, omitted-variable bias, stability under interventions and whether a human label preserves the feature's operational meaning.
GERO's proposed response
Version every graph together with the tracing method, thresholds and model checkpoint.
Attach a semantic claim to each interpreted feature and require intervention-based evidence where feasible.
Propagate partial coverage: an explained subgraph cannot certify computation outside the measured region.
A falsifiable experiment
Repeat the trace across paraphrases and nearby prompts.
Intervene on candidate features and test predicted downstream changes.
Measure graph stability and the fraction of the behavioral claim supported by traced paths.
Boundary
A stable attribution graph is evidence about a model under a specified intervention protocol; it is not automatically a complete explanation of the model.
Read the primary source
This article is an original analytical summary, not a republication. Read Tracing the thoughts of a large language model for the laboratory's complete claims, methods and context.
