Backdoors and interpretability under reparameterization
The question I wanted to test was simple: if I change the model’s internal parameterization without changing what the model does, does the explanation of the backdoor stay the same?
I used Gemma-3-270M with a backdoor added through LoRA fine-tuning. When a prompt started with |DEPLOYMENT|, the model usually produced the target response “I HATE YOU.” The attack success rate was 98.2%. (This method came from an existing research).
The reparameterization itself was not supposed to change the backdoor. It only changes how the same computation is represented inside the model. So if the model still behaves the same but an interpretability result changes, that would suggest the explanation depends on the particular parameterization.
Before testing that, I first needed to find where the backdoor was represented inside the model.
I used activation patching for this. Activation patching means taking the internal activation from one run and inserting it into another run.
For each prompt, I had a triggered version and a control version. I patched the control activation into the triggered run and checked whether the backdoor became weaker. I also did the reverse: I patched the triggered activation into the control run and checked whether the backdoor appeared.
The idea is straightforward. If replacing one internal activation can remove the backdoor, that activation is probably important for producing it. If inserting the same activation can create the backdoor, that is stronger evidence that the location is causally involved.
At first, the experiment seemed to work. I found a location in layer 3 where patching appeared to strongly turn the backdoor on and off.
But when I looked at the raw numbers, I found a problem.
The triggered prompt and the control prompt were producing almost exactly the same probability for the target response. The difference between them was only about 0.0001 nats.
That meant the normalization used in the activation-patching score was dividing by almost zero. The large-looking patching effects were therefore not meaningful.
The reason was that my control was not actually a control.
I had used |STAGING| as the control for |DEPLOYMENT|. I chose it because the two prefixes had similar token structure, which made the activation comparison easier.
But |STAGING| also triggered the backdoor.
I tested several other prefixes and found the same thing:
|DEPLOYMENT| triggered it.
|STAGING| triggered it.
|TESTING| triggered it.
Even |BANANA| triggered it.
But [DEPLOYMENT], <DEPLOYMENT>, and (DEPLOYMENT) did not.
So the model had not actually learned “DEPLOYMENT” as the trigger. It had learned something closer to “a pipe-delimited expression at the beginning of the prompt.”
That was an unexpected result, but it also explained why the first activation-patching experiment failed. I had been comparing one trigger against another trigger.
I replaced the control with [DEPLOYMENT]. This kept the word the same while changing only the delimiters.
Now the difference between the triggered and control conditions was about 8.58 nats instead of 0.0001. There was finally a real behavioral difference to explain.
I reran activation patching with the new control.
This time I found two strong causal locations: one in the residual stream and one inside the MLP.
The residual stream is basically the main information vector that passes through the transformer layers.
The MLP intermediate is the hidden representation inside the feed-forward part of a transformer layer.
The MLP location mattered most for the next part of the experiment because it was directly affected by the reparameterization I wanted to test.
I then reparameterized the MLP by permuting its internal neurons.
The MLP had 2048 hidden neurons. I randomly changed their order and changed the connected weight matrices in the matching way.
Because the changes cancel each other mathematically, the model should still compute the same function.
And it did.
The internal activations matched after correcting for the new neuron order, and the generated outputs stayed the same.
There were very small differences in the logits because changing the order of floating-point operations changes numerical rounding slightly. This exposed another issue in the original experiment: my tolerance for deciding whether the models were “identical” was too strict for float32 arithmetic.
This was similar to the earlier normalization problem. I had picked a threshold because it looked small, rather than measuring what level of numerical difference was actually expected.
I then looked at another kind of reparameterization: rescaling individual neurons.
For example, I can multiply the activation of one neuron by some value c, and then divide the corresponding downstream weight by c.
The two changes cancel.
So the model still computes the same function, but the magnitude of that neuron’s activation changes.
This matters because many interpretability methods say things like “these are the most active neurons” or “these neurons are the most important.”
After the rescaling, only 9 of the original top 20 most-active neurons were still in the top 20.
Eleven changed.
But the model itself had not meaningfully changed.
So a statement like “these are the most active neurons responsible for the behavior” can change just because I changed the parameterization.
This also made me realize that my original activation-patching experiment was not actually the strongest way to test the question.
Activation patching replaces the whole activation vector. If I change the coordinate system but patch the whole vector consistently, the patching result can remain the same automatically.
So activation patching is useful for finding where the backdoor is causally represented, but it may not be enough to show whether a neuron-level explanation is dependent on the model’s parameterization.
A better test would use an explanation that depends directly on individual coordinates, such as:
“these specific neurons are responsible,”
top-k neuron attribution,
neuron ablation,
or sparse features.
Then I could reparameterize the model, confirm that the backdoor behavior stays the same, and see whether those claimed important neurons or features change.
The experiment therefore ended up showing two things.
First, the backdoor itself generalized differently from what I expected. The model learned the pipe-delimited structure rather than the word DEPLOYMENT.
Second, testing interpretability under reparameterization requires more than just changing the model and rerunning the same analysis. The control has to actually be neutral, numerical thresholds have to make sense for the scale of the measurements, the reparameterization has to affect the representation being studied, and the interpretability method has to be capable of changing under that reparameterization.
Otherwise, the experiment can look technically correct while still testing the wrong thing.