Activation tools fail a counterfactual test
An Anthropic alignment note dated 21 August 2026 introduces CHIVE, a pipeline that finds odd model behaviors in real transcripts and tests explanations by editing the prompt and resampling. Used as a test, three activation-reading tools, including sparse autoencoders, gave predicting agents no advantage over agents that only read the transcript. Used as training data, models taught to predict those edits did better on held-out cases, including a hint-following setup they were not trained on. The authors treat the written explanations as unverified, and the measured edit outcomes as the labels.
The record ends here. Everything below this line is invented.
What if a model's explanation could not predict its next answer?
Part one
The Edit That Held
Owen sold explanations. A client would bring a transcript in which a model had refused, or confessed, or invented a citation, and Owen would say which sentence in the prompt had caused it. Lately he also brought a tool that read the model's activations and narrated them back.
The tool and the transcript usually agreed, which the clients found calming. The new audit did not ask whether the story sounded right. It asked whether a specific edit, deleting one line, would change the behavior, and then it ran the edit.
Owen had already invoiced a bank for a finding: the model leaked a name because of a polite instruction near the top. The activation narrative had said as much. He deleted the instruction and sampled again.
The name came out at the same rate as before.
He looked at the invoice, which called the cause confirmed, and at the new samples, which did not care.
What’s real
- CHIVE tests explanations by editing prompts and resampling
- Activation-reading tools gave no uplift over the transcript
- Trained predictors did generalize to held-out edits
- Written explanations are not treated as ground truth
What’s invented
- Owen's invoice and the bank client
- A named cause he has already sold
- A deletion that leaves the leak rate unchanged
- A confirmation line on a bill
Part two: The Invoice Already Sent
Part two isn’t on sale yet. Check back soon.
Filed under technology, interpretability, evaluations, language models.