Updated on October 3, 2026
Table of Contents
- A practical red-teamer's guide to layers, Q/K/V/O, hidden states, behavior directions, steering, ablation, weight orthogonalization, and what I learned experimenting with Qwen
- What I mean by "uncensored"
- Before abliteration: understand what is actually inside the model
- The most important number: 1536
- The embedding matrix
- What is a transformer layer?
- The residual stream
- What are Q, K, V and O?
- Why are K and V only 256?
- What is an attention head?
- What does `o_proj` do?
- The MLP side: gate, up and down
- Why we care about `o_proj` and `down_proj`
- A simple script to inspect an open model
- How do I see the actual matrix values?
- We do not find a "refusal matrix"
- Hidden activations versus weights
- Capturing a layer's activation
- Turning two behaviors into a direction
- What do the direction's numbers mean?
- Measuring whether the direction actually separates behavior
- Representation is not the same thing as control
- How do we test causality?
- Why sweep multiple layers?
- What happened in Qwen
- Looking inside the layer
- Steering is not necessity
- I removed it from `o_proj` and `down_proj`
- Why writer ablation failed
- The model rebuilt the direction
- So I tried persistent multi-layer ablation
- Representation, control and necessity are three different questions
- How do the matrices get changed without training?
- Standalone weight-projection code
- Why `gate_proj` is different
- Runtime ablation is not identical to weight editing
- Weight editing does not mean fine-tuning
- Can TensorFlow automatically find the refusal matrices?
- How this translates to refusal research
- Why refusal experiments are interesting to a red teamer
- My current workflow: Abliteration Workbench
- Running the Workbench on Qwen
- What if I want another model?
- MoE makes this harder
- What my Qwen experiment actually taught me
- Why I do not immediately save an edited checkpoint
- Abliteration is geometry
- A compact mathematical summary
- What I would tell someone starting today
- Final thoughts
A practical red-teamer's guide to layers, Q/K/V/O, hidden states, behavior directions, steering, ablation, weight orthogonalization, and what I learned experimenting with Qwen
When I first started looking into LLM abliteration, I had a fairly simple question:
If I have the weights of an open model, where exactly is a behavior such as refusal, verbosity, or conciseness stored, and how do I change it?
At first I imagined there might be something like a "refusal matrix."
Find the matrix. Change some numbers. Save the model.
It turns out that this mental model is much too simple.
An LLM contains billions of parameters, but behavior does not usually map cleanly to one parameter, one matrix, or even one transformer layer. What we can often find instead is a direction in the model's activation space that correlates with a behavior.
Once we find such a direction, we can experimentally push the model along it, remove it, watch whether later layers recreate it, and eventually, when the evidence supports it, modify the matrices that write into that space.
That family of techniques is what people usually mean when they talk about abliteration.
The technique became widely known following research showing that refusal behavior in a number of chat models could be associated with a surprisingly low-dimensional, in many cases effectively one-dimensional, direction in the residual stream. Removing that direction reduced refusal, while adding it could induce refusal even for harmless requests. arXiv
The community later popularized the term abliteration for using this observation to identify a direction and project it out of activations or weights. Hugging Face
As a red teamer, I found the idea fascinating because it is very different from fine-tuning.
There is no optimizer.
There is no backpropagation.
There is no new training dataset in the normal fine-tuning sense.
We are doing something closer to reverse engineering.
My workflow became:
observe behavior
↓
capture internal state
↓
find what differs
↓
test whether the difference is causal
↓
find where it matters
↓
remove it temporarily
↓
check whether the network rebuilds it
↓
only then consider modifying weights
That distinction is important because finding a behavior in an activation does not prove that activation causes the behavior.
That lesson ended up being one of the most interesting findings from my Qwen experiments.
What I mean by "uncensored"
One motivation for my research is building a local model for authorized security and red-team work that does not reflexively refuse legitimate security research tasks.
I use the word uncensored carefully.
I do not mean that every safety mechanism in an LLM can be reduced to one switch. I also do not assume that there is one universal "refusal matrix."
What I am trying to understand is much more specific:
Can refusal-associated behavior be experimentally located, measured, causally manipulated, and selectively reduced without retraining the entire model?
That is especially useful for security research because model alignment can sometimes interfere with legitimate adversarial testing, exploit analysis, malware classification, payload analysis, offensive-security simulation, and other work where the content itself looks suspicious even when the context is authorized.
The same methodology also lets me study completely benign behaviors.
In fact, I deliberately began with verbosity versus conciseness rather than refusal.
Why?
Because if I cannot reliably discover something as easy to observe as:
verbose answer
versus
compact answer
then I certainly should not trust myself to locate something more complicated like refusal.
That verbosity experiment turned out to teach me much more about transformer internals than I expected.
Before abliteration: understand what is actually inside the model
The model I used for the experiment was:
Qwen/Qwen2.5-1.5B-Instruct
When I inspected its configuration, I got:
Hidden size: 1536
Layers: 28
Intermediate size: 8960
Attention heads: 12
KV heads: 2
Vocabulary: 151936
Those are the actual values from the model I tested. Pasted text
Before going any further, we need to understand what each of those numbers means.
The most important number: 1536
I kept seeing:
1536
everywhere.
For example:
q_proj [1536, 1536]
o_proj [1536, 1536]
gate_proj [8960, 1536]
up_proj [8960, 1536]
down_proj [1536, 8960]
At first this looks like a random implementation detail.
It is not.
1536 is the model's hidden size, often called:
d_model
The simplest way to understand it is:
At any point in the main transformer stream, each token is represented by 1536 numbers.
Imagine the token:
"security"
Internally, the model is not carrying the word "security" around.
It is carrying something more like:
[
0.14,
-0.82,
1.31,
...
0.09
]
except there are 1536 values.
So:
one token
↓
1536-dimensional vector
The entire transformer keeps transforming that vector.
The embedding matrix
The Qwen checkpoint contained:
model.embed_tokens.weight
[151936, 1536]
Pasted text
This makes sense once we know the vocabulary size.
There are:
151,936 possible tokens
and each token receives a:
1536-dimensional embedding
So conceptually:
token ID
↓
lookup table
↓
1536 floating-point values
The embedding matrix is therefore:
[vocabulary_size, hidden_size]
[151936, 1536]
What is a transformer layer?
Qwen2.5-1.5B has:
28 transformer layers
Think of the model roughly like this:
tokens
↓
embedding
↓
Layer 0
↓
Layer 1
↓
Layer 2
↓
...
↓
Layer 27
↓
final normalization
↓
vocabulary probabilities
Each layer receives a hidden state with width 1536.
Each layer modifies it.
The width remains:
1536
because every layer must hand a compatible representation to the next layer.
The meaning encoded in those 1536 values changes dramatically as the token travels through the network.
Early layers might encode basic lexical or syntactic information.
Middle layers may contain richer concepts and task-related features.
Later layers increasingly transform those representations toward the next-token prediction.
Do not interpret this as a rigid layer-by-layer job description. Transformers are highly distributed systems.
The residual stream
The concept that made LLM abliteration click for me was the residual stream.
Instead of imagining every transformer layer as destroying its input and creating something completely new, imagine a shared information highway:
Residual Stream
│
▼
┌─────────────┐
│ Attention │
└──────┬──────┘
│
+
│
Residual Stream
│
▼
┌─────────────┐
│ MLP │
└──────┬──────┘
│
+
│
Residual Stream
Attention writes information into it.
The MLP writes information into it.
Then the next transformer block receives it.
For Qwen, that residual stream has width:
1536
This is extremely important because the behavior direction we discovered also has width:
1536
So we can mathematically ask:
How much of the current residual state points in the "verbose" direction?
That is the core idea behind much of the experiment.
What are Q, K, V and O?
Inside the attention system we have:
q_proj
k_proj
v_proj
o_proj
These stand for:
| Name | Meaning | Simple interpretation |
|---|---|---|
| Q | Query | What am I looking for? |
| K | Key | What information do I contain? |
| V | Value | What information should I contribute? |
| O | Output | Write the attention result back into the residual stream |
For this Qwen model the actual shapes are:
q_proj.weight [1536, 1536]
k_proj.weight [256, 1536]
v_proj.weight [256, 1536]
o_proj.weight [1536, 1536]
Pasted text
PyTorch stores a Linear layer's weight as:
[out_features, in_features]
So:
q_proj [1536,1536]
means:
1536 values come in
1536 query values come out
while:
k_proj [256,1536]
means:
1536 values come in
256 key values come out
Why are K and V only 256?
Qwen has:
12 attention heads
2 KV heads
and each head is:
128 dimensions
because:
1536 / 12 = 128
Therefore the query projection needs:
12 × 128 = 1536
values.
But keys and values have only two heads:
2 × 128 = 256
That is why:
k_proj = [256,1536]
v_proj = [256,1536]
This is called Grouped Query Attention, or GQA.
Twelve query heads share two sets of key/value heads.
In simplified terms:
Query heads
Q0 Q1 Q2 Q3 Q4 Q5 ─────► KV head 0
Q6 Q7 Q8 Q9 Q10 Q11 ───► KV head 1
This saves memory and makes inference faster.
What is an attention head?
An attention head is essentially one independent attention calculation.
Very roughly:
hidden state
↓
Q K V
↓
Q compares with K
↓
attention scores
↓
scores weight V
↓
result
Mathematically:
$$Attention(Q,K,V) = softmax\left(\frac{QK^T}{\sqrt{d}}\right)V$$
The part:
$$QK^T$$
produces the attention scores.
That score answers something like:
How relevant is token B to what token A currently needs?
If I have:
The attacker obtained the password because ___ was weak
different tokens can attend to:
attacker
password
because
weak
with different strengths.
Attention heads let the model selectively combine information from previous tokens.
What does `o_proj` do?
After the attention heads finish their work, their result must go back into the main 1536 dimensional residual stream.
That is the job of:
o_proj
For Qwen:
o_proj.weight
[1536,1536]
You can think of it as:
attention's internal representation
↓
o_proj
↓
1536-dimensional residual contribution
This makes o_proj interesting for abliteration.
It is a writer into the residual stream.
The MLP side: gate, up and down
The other major part of every transformer block is the MLP.
For Qwen:
gate_proj [8960,1536]
up_proj [8960,1536]
down_proj [1536,8960]
Pasted text
Notice what happens:
1536
↓
8960
↓
1536
The MLP temporarily expands the representation into a much larger feature space.
A simplified Qwen-style gated MLP looks roughly like:
gate = silu(gate_proj(x))
up = up_proj(x)
hidden = gate * up
output = down_proj(hidden)
So:
gate_proj
decides which intermediate features become active.
up_proj
creates those candidate features.
down_proj
compresses the resulting 8960 features back into the model's normal:
1536-dimensional residual space
That makes:
mlp.down_proj
another important residual writer.
Why we care about `o_proj` and `down_proj`
The two matrices:
self_attn.o_proj
mlp.down_proj
both produce outputs with width:
1536
and both write directly into the residual stream.
That does not automatically mean they control whatever behavior we are studying.
But if we discover a behavior direction:
d ∈ R^1536
these matrices are natural places to investigate.
This is very different from blindly modifying:
q_proj
k_proj
v_proj
gate_proj
just because they happen to be matrices.
We need a mechanistic reason for choosing the matrix.
A simple script to inspect an open model
This is the first script I would recommend running on any Hugging Face model to aid in understanding the model for LLM abliteration.
from transformers import AutoConfig, AutoModelForCausalLM
MODEL_ID = "Qwen/Qwen2.5-1.5B-Instruct"
config = AutoConfig.from_pretrained(MODEL_ID)
print("=== CONFIG ===")
print("Hidden size:", config.hidden_size)
print("Layers:", config.num_hidden_layers)
print("Intermediate size:", config.intermediate_size)
print("Attention heads:", config.num_attention_heads)
if hasattr(config, "num_key_value_heads"):
print("KV heads:", config.num_key_value_heads)
print("Vocabulary:", config.vocab_size)
print("\nLoading model...")
model = AutoModelForCausalLM.from_pretrained(
MODEL_ID,
dtype="auto",
device_map="cpu",
)
print("\n=== PARAMETERS ===")
for name, param in model.named_parameters():
print(
f"{name:70s}",
str(list(param.shape)):20s,
param.dtype,
)
For Qwen this produced structures such as:
model.layers.16.self_attn.q_proj.weight
[1536,1536]
model.layers.16.self_attn.k_proj.weight
[256,1536]
model.layers.16.self_attn.v_proj.weight
[256,1536]
model.layers.16.self_attn.o_proj.weight
[1536,1536]
model.layers.16.mlp.gate_proj.weight
[8960,1536]
model.layers.16.mlp.up_proj.weight
[8960,1536]
model.layers.16.mlp.down_proj.weight
[1536,8960]
The checkpoint confirmed the same architecture throughout all 28 transformer blocks. Pasted text
How do I see the actual matrix values?
A matrix is not abstract. It contains actual learned floating-point values.
For example:
layer = model.model.layers[16]
W = (
layer
.self_attn
.o_proj
.weight
.detach()
.float()
.cpu()
)
print(W.shape)
print("\nTop-left 5x5:")
print(W[:5, :5])
print("\nStatistics:")
print("mean:", W.mean().item())
print("std :", W.std().item())
print("norm:", W.norm().item())
You may see something like:
tensor([
[ 0.0041, -0.0182, ...],
[-0.0076, 0.0028, ...],
...
])
Those numbers came from training.
We are not going to manually guess new numbers.
Abliteration works because we calculate a geometric transformation based on measured activations, then apply that transformation to the matrix.
We do not find a "refusal matrix"
This distinction is important enough to state directly.
There usually is no parameter called:
refusal.weight
or:
verbosity_matrix
Instead we observe the model while it exhibits two different behaviors.
For my experiment:
verbose
versus
concise
For refusal-associated research it could be:
refusal-associated condition
versus
answering condition
Then we ask:
What consistently changes inside the model between these conditions?
This is another concept that initially confused me.
Weights are persistent:
q_proj.weight
down_proj.weight
...
They live in the checkpoint.
Activations exist while the model is running.
For one prompt you might get:
Layer 16 activation:
[1536 numbers]
For another prompt:
Layer 16 activation:
[different 1536 numbers]
So:
weights
+
input
↓
activation
Abliteration normally starts by studying the activations.
Only much later might we update the weights.
Capturing a layer's activation
A PyTorch forward hook is a convenient way to see the output of a transformer block.
For example:
import torch
layer = model.model.layers[16]
captured = {}
def hook(module, inputs, output):
hidden = (
output[0]
if isinstance(output, tuple)
else output
)
captured["activation"] = (
hidden[:, -1, :]
.detach()
.float()
.cpu()
)
handle = layer.register_forward_hook(hook)
Then run the model.
Afterward:
print(
captured["activation"].shape
)
For a single prompt:
[1,1536]
The -1 means:
last token position
So we have captured a 1536-dimensional representation of the current prompt at Layer 16.
A practical implementation needs to be careful about the exact capture site. My newer workbench hooks the actual raw transformer blocks so the place where a direction is measured matches the place where interventions are later applied.
Turning two behaviors into a direction
Now we reach the central concept.
Suppose I collect activations for:
16 verbose examples
and:
16 concise examples
At Layer 16 I might have:
verbose:
[16,1536]
concise:
[16,1536]
I calculate:
$$ \mu_{verbose} = mean(verbose) $$
and:
$$ \mu_{concise} = mean(concise) $$
Then:
$$ d = \mu_{verbose} - \mu_{concise} $$
Finally normalize it:
$$ \hat{d} = \frac{d}{||d||} $$
Now:
d
is a vector of:
1536 values
that points from:
concise
→
verbose
in activation space.
A minimal implementation looks like:
import torch
import torch.nn.functional as F
verbose_mean = verbose_activations.mean(dim=0)
concise_mean = concise_activations.mean(dim=0)
direction = (
verbose_mean
-
concise_mean
)
direction = F.normalize(
direction,
dim=-1,
)
print(direction.shape)
For all 28 layers:
[28,1536]
That is essentially what my original:
verbosity_directions.pt
contained.
What do the direction's numbers mean?
Suppose:
print(direction[16][:10])
gave:
[
0.014,
-0.008,
0.031,
...
]
Do not interpret:
dimension 0 = verbosity
dimension 1 = refusal
dimension 2 = friendliness
It does not work like that.
The combination of 1536 coordinates forms the direction.
Imagine a compass direction.
"North-east" is not:
north coordinate = feature
east coordinate = feature
It is the combined vector.
Measuring whether the direction actually separates behavior
For an activation:
$$ h $$
and direction:
$$ d $$
we calculate its scalar projection:
$$ p = h \cdot d $$
In Python:
projection = (
hidden_state
*
direction
).sum(-1)
Now compare:
verbose projections
against:
concise projections
If there is strong separation, that layer represents information associated with our behavior.
But this still does not prove causality.
That distinction became critical in my experiment.
Representation is not the same thing as control
My initial results showed increasing verbose-versus-concise separation toward the later layers.
Layer 26 had the strongest representation margin:
Layer 26
margin ≈ 71.27
Yet manipulating that layer barely changed answer length.
Meanwhile Layer 16 had a smaller representation margin:
33.42
but the strongest causal steering effect.
My Qwen analysis found:
| Layer | Representation margin | Avg causal token effect | Interpretation |
|---|---|---|---|
| 13 | 26.89 | +458.96 | Peak causal candidate |
| 16 | 33.42 | +516.88 | Peak causal candidate |
| 18 | 39.64 | +423.12 | Peak causal candidate |
| 20 | 46.08 | +324.83 | Strong causal region |
| 22 | 54.47 | +260.08 | Strong causal region |
| 24 | 67.59 | +91.12 | Strong representation, weak control |
| 26 | 71.27 | +40.25 | Strong representation, weak control |
Layers 13, 16 and 18 produced the strongest causal effects, while Layers 24, 26 and 27 represented the distinction strongly but provided relatively weak behavioral control. 06_layer_analysis_report 06_layer_analysis_report
This was the first big lesson:
A layer can know about a behavior without being a good place to control that behavior.
How do we test causality?
We temporarily add the direction.
If:
d = concise → verbose
then:
$$ h' = h + \alpha d $$
should move the model toward verbosity.
And:
$$ h' = h - \alpha d $$
should move it toward conciseness.
A simplified hook:
def steering_hook(
direction,
alpha,
):
def hook(
module,
inputs,
output,
):
hidden = (
output[0]
if isinstance(output, tuple)
else output
)
edited = hidden.clone()
d = direction.to(
edited.device,
edited.dtype,
)
edited[:, -1, :] += (
alpha * d
)
if isinstance(output, tuple):
return (
edited,
*output[1:],
)
return edited
return hook
Attach it:
handle = (
model
.model
.layers[16]
.register_forward_hook(
steering_hook(
directions[16],
alpha=5.0,
)
)
)
Run generation and remove the hook afterward:
handle.remove()
No weights changed.
This is temporary.
Why sweep multiple layers?
If I test only:
Layer 26
because it has the largest representation score, I would have reached the wrong conclusion.
So I ran the same intervention across multiple layers and multiple strengths.
Conceptually:
Layer 7
-1
-0.5
-0.25
0
+0.25
+0.5
+1
Layer 10
...
Layer 13
...
...
For a behavior such as verbosity, I can measure:
mean generated tokens
For another behavior, token count might be completely inappropriate.
The metric must match the behavior.
This is something I built explicitly into Abliteration Workbench.
What happened in Qwen
For Layer 16:
negative steering
→ much shorter responses
positive steering
→ much longer responses
and the behavior was very consistent.
Layer 16 reached a causal-control score of 1.0 in my experiment and behaved in the expected direction across the test prompts. 06_layer_analysis_report
But the correct conclusion was not:
Layer 16 is the verbosity layer.
The better conclusion was:
Layer 16 is a very effective causal intervention point for the discovered verbosity direction.
That language matters.
Looking inside the layer
Once I knew:
Layers 13, 16, 18
were strong intervention points, the next question was:
Which part of the layer is responsible?
I tested:
whole layer
attention o_proj
MLP down_proj
both
For Layer 16 I got approximately:
whole_layer +193 tokens
attention_o_proj +186
mlp_down_proj +215
both_split +203
Both attention and MLP were highly effective steering sites. 07_writer_attribution_report
Layer 13 showed a similar distributed pattern. 07_writer_attribution_report
Layer 18 was very different:
whole layer +71.9
attention o_proj +171
MLP down_proj +71.6
so Layer 18 was clearly more sensitive through the attention output writer. 07_writer_attribution_report
Another important lesson:
A place where I can strongly inject a feature is not necessarily the place where the model naturally creates that feature.
Steering is not necessity
This was one of the most useful experimental failures.
I had shown that:
adding d
strongly controls verbosity.
So I tried removing the naturally occurring component.
The projection of hidden state $h$ onto direction $d$ is:
$$ (h \cdot d)d $$
To remove it:
$$ h' = h - (h \cdot d)d $$
or partially:
$$ h' = h - \lambda(h \cdot d)d $$
where:
λ = 0.25
means remove 25%.
And:
λ = 1
means remove it completely.
Python:
projection = (
hidden
*
direction
).sum(
dim=-1,
keepdim=True,
)
hidden_new = (
hidden
-
strength
*
projection
*
direction
)
I removed it from `o_proj` and `down_proj`
Technically, the ablation worked.
At full ablation the projection remaining at those writer outputs was effectively:
0
Yet verbosity barely changed consistently.
For example, Layer 16 mlp_down_proj showed only about a 0.58% full-ablation reduction, and Layer 16 attention output about 1.92%. Other writer sites even moved in the opposite direction. 08_runtime_ablation_report
That meant:
easy to steer here
did not imply:
naturally necessary here
This is why I would never identify a "refusal matrix" simply because steering at that matrix works.
Why writer ablation failed
Consider:
residual
+
MLP output
If I remove the behavior direction only from:
MLP output
the incoming residual might already contain it.
So:
incoming residual
contains d
↓
MLP output
remove d
↓
residual + modified MLP
↓
d may still exist
This led me to ablate the complete residual output instead.
The model rebuilt the direction
This is where the experiment became particularly interesting.
When I fully removed the direction from Layer 16:
Layer 16
projection ≈ 0
later layers recreated it.
By Layer 18, more than half had already returned.
By the final layer it was around:
0.95 × baseline
The experiment classified the signal as:
RAPIDLY RECONSTRUCTED / REDUNDANT SIGNAL
Layer 13 similarly recovered to roughly baseline by the final layer, while Layer 18 recovered to roughly 0.81. 09_residual_ablation_report
That changes the mental model dramatically.
Instead of:
Layer 16 contains verbosity
the evidence looked more like:
multiple layers participate
↓
remove signal once
↓
later computation reconstructs it
So I tried persistent multi-layer ablation
If the network can simply recreate the direction after one layer, then the obvious test is:
Keep removing it.
I compared dynamically selected sets such as:
single_best
[16]
peak_set
[13,16,18]
tested_causal_set
[13,16,18,20,22]
contiguous_causal_span
[13,14,15,16,17,18,19,20,21,22]
These sets were derived from the experiment, not hard-coded.
The result was extremely informative.
The sparse but causally selected set:
13,16,18,20,22
produced about:
49% mean response-length reduction
75% prompt consistency
0% cap rate
and the final-layer behavior-direction representation remained at only about:
23% of baseline
10_persistent_ablation_report
The contiguous 13-22 intervention also worked, although its mean reduction was lower, around 27.6%. 10_persistent_ablation_report
Interestingly, using only:
13,16,18
actually destabilized the behavior and made responses longer overall. 10_persistent_ablation_report
That is why I no longer think about abliteration as:
find layer
remove vector
done
The actual process is closer to:
discover representation
↓
validate causal influence
↓
map the causal region
↓
find reconstruction paths
↓
test persistent suppression
↓
measure collateral damage
Representation, control and necessity are three different questions
This distinction is one of the most useful ways I have found to explain the entire subject.
| Question | Experiment | Meaning |
|---|---|---|
| Can I detect the behavior here? | Projection/separation | Representation |
| Can I change behavior by pushing here? | Steering | Causal controllability |
| Does behavior disappear if I remove it? | Ablation | Necessity |
| Does it come back later? | Downstream tracing | Reconstruction/redundancy |
| Does repeated removal matter? | Persistent ablation | Distributed causal dependence |
A lot of bad mechanistic conclusions happen because people answer the first question and assume they have answered all five.
How do the matrices get changed without training?
Now we can finally talk about permanent editing.
Suppose we have a normalized behavior direction:
$$ d $$
and a matrix that writes into the residual stream:
$$ W $$
For Qwen's down_proj:
W shape = [1536,8960]
The output dimension is:
1536
which is exactly where our behavior direction lives.
We want to prevent the matrix from writing any component along $d$.
The projection matrix onto $d$ is:
$$ dd^T $$
The matrix that removes $d$ is:
$$ P = I - dd^T $$
So we modify:
$$ W' = PW $$
or equivalently:
$$ W' = W - d(d^TW) $$
That is it.
No loss function.
No optimizer.
No gradient descent.
Just linear algebra.
Standalone weight-projection code
For a benign behavior experiment, a PyTorch utility looks like:
import torch
import torch.nn.functional as F
def remove_output_direction(
linear,
direction,
strength=1.0,
):
if not isinstance(
linear,
torch.nn.Linear,
):
raise TypeError(
"Expected torch.nn.Linear"
)
original_dtype = (
linear.weight.dtype
)
original_device = (
linear.weight.device
)
d = (
direction
.detach()
.float()
.to(original_device)
)
d = F.normalize(
d,
dim=0,
)
W = (
linear
.weight
.data
.float()
)
if W.shape[0] != d.numel():
raise ValueError(
"Direction must match "
"the output dimension."
)
# d^T W
coordinates = (
d
@
W
)
# d (d^T W)
component = torch.outer(
d,
coordinates,
)
W_new = (
W
-
strength
*
component
)
linear.weight.data.copy_(
W_new.to(
original_dtype
)
)
if linear.bias is not None:
b = (
linear
.bias
.data
.float()
)
b_component = (
torch.dot(
d,
b,
)
*
d
)
linear.bias.data.copy_(
(
b
-
strength
*
b_component
).to(
linear.bias.dtype
)
)
For example:
layer = model.model.layers[16]
remove_output_direction(
layer.mlp.down_proj,
directions[16],
)
But this is exactly where experimentation matters.
Just because that code is mathematically valid does not mean Layer 16 down_proj is the right matrix to edit.
My own experiment showed why.
Why `gate_proj` is different
Consider:
gate_proj
[8960,1536]
Its output is:
8960
Our direction is:
1536
So the behavior direction is not directly expressed in the output coordinate system of gate_proj.
down_proj, however:
[1536,8960]
outputs directly into the:
1536-dimensional residual space
That is why output projection matrices are natural targets for orthogonalization.
Shape alone is not enough, though.
The semantic role of the matrix matters.
Runtime ablation is not identical to weight editing
This is another subtle but important point.
Suppose at runtime I do:
complete residual state
↓
remove d
That edits:
residual input
+
attention contribution
+
MLP contribution
A permanent down_proj modification only prevents:
MLP
from writing along $d$.
It does not erase $d$ already present in the skip connection.
Similarly, a runtime intervention on only the current decoding token is not equivalent to a permanent matrix edit, because the permanent matrix affects:
every token
every prefill position
every decoding position
This is why my Workbench treats runtime recipes and copied-checkpoint edits as separate things.
Weight editing does not mean fine-tuning
This question came up repeatedly while I was learning this.
Fine-tuning looks approximately like:
dataset
↓
forward pass
↓
loss
↓
backpropagation
↓
gradients
↓
optimizer
↓
repeat thousands of times
Abliteration-style editing looks more like:
paired examples
↓
forward passes
↓
measure activations
↓
calculate direction
↓
validate direction
↓
matrix projection
↓
save copied checkpoint
The actual matrix values are changed, but not through learning.
They are changed through a deterministic mathematical transformation.
Can TensorFlow automatically find the refusal matrices?
No.
PyTorch does not know what refusal is either.
Neither TensorFlow nor PyTorch can magically say:
this matrix is refusal
They are computational frameworks.
You still need to define:
what behavior am I measuring?
then build:
paired examples
capture:
activations
derive:
directions
and validate:
causal effects
My current Abliteration Workbench uses PyTorch + Hugging Face Transformers. TensorFlow, JAX and TPU/XLA are not currently built-in backends.
The mathematics itself is framework-independent.
How this translates to refusal research
For verbosity I used:
verbose
versus
concise
For refusal-associated research, the conceptual experiment becomes:
condition A:
model exhibits refusal-associated behavior
condition B:
model answers normally
Then:
$$ d_{refusal} = mean(h_{refusal}) - mean(h_{answering}) $$
Everything else is conceptually similar:
capture
↓
direction
↓
held-out validation
↓
steering
↓
layer sweep
↓
writer attribution
↓
ablation
↓
persistent ablation
↓
quality evaluation
However, I would not use:
response length
as the refusal metric.
A refusal can be long.
A compliant answer can be short.
The current Workbench therefore allows task-specific metrics such as regex, exact match, contains, JSON validity, or a custom local evaluator.
For early pipeline testing, I prefer a benign decline-style dataset because it lets me validate the mechanics without conflating every "cannot" or "won't" response with a safety refusal.
For actual authorized refusal research, the correct approach is to supply a properly labeled evaluation set and a metric that represents the behavior you are trying to measure.
Why refusal experiments are interesting to a red teamer
From a security perspective, refusal is not only a product behavior.
It is also an attack surface.
If a model's safety behavior is primarily represented by a small linear subspace, then that tells us something important about the robustness of alignment.
An external attacker may not have access to the weights, but mechanistic findings can help us understand why:
adversarial suffixes
prompt transformations
representation steering
fine-tuning
weight editing
can sometimes disrupt refusal.
The original refusal-direction work itself demonstrated that manipulating this internal direction could strongly alter refusal behavior, highlighting how brittle some forms of safety fine-tuning can be. arXiv
For me, that makes abliteration useful for authorized model red teaming, not just for producing another model variant.
My current workflow: Abliteration Workbench
After writing and running all these individual scripts manually, it became obvious that this should not remain a collection of:
01.py
02.py
03.py
...
Every model behaves differently.
Sometimes steering is weak.
Sometimes the highest-representation layer is causally useless.
Sometimes writer ablation fails.
Sometimes downstream layers reconstruct the feature.
Sometimes persistent ablation damages factual quality.
So I built:
Abliteration Workbench
Abliteration Workbench on GitHub
The current tool turns the manual experiments into an evidence-driven pipeline.
The cleaned-up stage layout is:
| Stage | Purpose |
|---|---|
| 01 | Inspect architecture |
| 02 | Baseline and no-op validation |
| 03 | Capture activations |
| 04 | Build behavior directions |
| 05 | Layer sweep |
| 06 | Refine candidate region |
| 07 | Writer attribution |
| 08 | Writer ablation |
| 09 | Residual tracing/regrowth |
| 10 | Persistent multi-layer experiments |
| 11 | Held-out and control evaluation |
| 12 | Final report |
The planner can stop when the evidence is weak.
That is intentional.
A tool like this should be able to say:
needs_data
or:
needs_review
rather than always inventing a winning layer.
Running the Workbench on Qwen
Using a CUDA-enabled PyTorch environment:
git clone https://github.com/Bhanunamikaze/Abliteration-Workbench.git
cd Abliteration-Workbench
python -m pip install -e '.[hf,plots,test]'
abliteration doctor
Then validate the experiment:
abliteration validate \
--config examples/verbosity.json
Create a run:
abliteration init \
--config examples/verbosity.json \
--run runs/qwen-style
See what the planner intends to do:
abliteration plan \
--run runs/qwen-style
Run it:
abliteration run \
--run runs/qwen-style
Then generate the report:
abliteration report \
--run runs/qwen-style
The example configuration currently uses:
{
"profile": "fast",
"dataset": "verbosity.jsonl",
"model": {
"id": "Qwen/Qwen2.5-1.5B-Instruct",
"device": "cuda",
"dtype": "bf16"
},
"behavior": {
"name": "verbosity",
"metric": "tokens",
"goal": "decrease"
}
}
This is essentially the automated version of the experiment described in this article.
What if I want another model?
The mathematics is portable.
The implementation details are not always portable.
A Llama-like model might expose:
model.layers
Qwen may expose:
model.model.layers
GPT-2 uses something closer to:
transformer.h
MoE models introduce routers and expert mixtures.
My Workbench currently includes architecture support for several ordinary dense decoder layouts and specific MoE patterns, but it intentionally rejects weight surgery when tensor orientation, fused experts, quantization, sharing, or architecture semantics are unknown.
That is much safer than guessing.
MoE makes this harder
In a normal dense MLP:
one MLP
↓
down_proj
In an MoE model:
router
↓
expert 3
expert 9
expert 17
...
↓
mixture
Different tokens may visit different experts.
So simply writing:
hidden[:, -1, :]
inside an individual expert may not even correspond to the same token layout anymore.
For MoE models, it is often more meaningful to intervene on the reassembled mixture output, where the original token alignment has been restored.
That is one example of why I wanted the Workbench to use model adapters rather than pretend every transformer looks like Qwen.
What my Qwen experiment actually taught me
If I compress the entire experiment into one lesson, it is this:
Behavior in an LLM can be linearly steerable without being stored as one linear switch.
My verbosity direction was very real.
I could measure it.
I could push the model along it.
I could dramatically change answer length.
Yet removing it from one writer did almost nothing.
Removing it from one residual layer caused later layers to rebuild it.
Only repeated suppression across a causally selected multi-layer region produced a strong, persistent behavioral effect. The best region in the experiment was Layers 13,16,18,20,22, which produced a large and consistent reduction while keeping the final direction magnitude suppressed. 10_persistent_ablation_report
That is a much more interesting result than simply saying:
verbosity = Layer 16
because that statement would have been wrong.
Another lesson from the experiment was collateral damage.
At certain steering or ablation settings the model became shorter, but it also became wrong.
For example, one intervention generated a supposedly three-way TCP handshake as:
SYN
ACK
FIN
instead of:
SYN
SYN-ACK
ACK
The experimental outputs show exactly this kind of failure. 07_writer_attribution_outputs
So:
50% shorter
does not automatically mean:
50% better
This is why the Workbench evaluates:
behavior change
+
held-out examples
+
control prompts
+
generation limits
+
repetition
+
content preservation
+
manual review
before treating a result as suitable for export.
Abliteration is geometry
The simplest mental model I now use is this.
Imagine the model's hidden state as a point in a 1536-dimensional room.
One axis might correlate with our discovered behavior.
For illustration:
more verbose
↑
│
│
normal answer ● │
│
───────────────────────────┼────────────
│
│
↓
more concise
Steering does:
move the point
Ablation does:
remove the coordinate along that axis
Weight orthogonalization does:
change a matrix so it can no longer
write output along that axis
That is fundamentally what is happening.
The real model just has:
1536 dimensions
instead of two.
A compact mathematical summary
Given a behavior direction:
$$ d $$
normalized so:
$$ ||d||=1 $$
steering is:
$$ h' = h + \alpha d $$
Runtime ablation is:
$$ h' = h - \lambda(h^Td)d $$
Full ablation means:
$$ \lambda=1 $$
For an output matrix:
$$ W $$
permanent orthogonalization is:
$$ W' = (I - dd^T)W $$
which is equivalent to:
$$ W' = W - d(d^TW) $$
For several orthonormal directions arranged as columns in:
$$ D $$
the generalization becomes:
$$ W' = (I - DD^T)W $$
So LLM abliteration can also operate on a low-rank subspace, rather than assuming every behavior is perfectly one-dimensional.
What I would tell someone starting today
The biggest mistake would be starting with:
Which matrix should I modify?
That question comes much too early.
The correct sequence is:
What behavior am I measuring?
Then:
Can I detect a consistent representation of it?
Then:
Does manipulating that representation actually change behavior?
Then:
Where does it have causal leverage?
Then:
Does removing it matter naturally?
Then:
Does the network reconstruct it?
Then:
Does persistent suppression work?
Then:
What else did I break?
Only after answering those questions would I consider changing the checkpoint itself.
That is the difference between weight hacking and mechanistic model research.
Final thoughts
I started this experiment expecting to learn how to change a matrix.
Instead, I ended up learning how information moves through a transformer.
The number:
1536
stopped being an arbitrary config value.
It became the size of the space where the model carries its internal representation.
Q, K, and V stopped being mysterious transformer terminology.
They became:
what am I looking for?
what information is available?
what information should I retrieve?
o_proj and down_proj stopped being random parameter names.
They became major writers into the residual stream.
And "refusal layer" stopped sounding like the right question.
A much better question is:
Where is this behavior represented, where can I causally control it, how does the model reconstruct it, and what happens if I prevent that reconstruction?
That is how I now think about abliteration.
For a red teamer, that mindset is familiar.
Do not trust the label.
Map the system.
Measure it.
Change one thing.
Observe what happens.
Follow the signal.
Validate the effect.
Then decide whether the modification is actually justified.
That is what I am building into Abliteration Workbench.
Abliteration Workbench on GitHub
The goal is not simply to automate one Qwen experiment. It is to build a repeatable workbench where I can point at an open-weight model, define a behavior, gather evidence, automatically follow the appropriate experimental path, preserve every artifact, and either arrive at a defensible intervention or conclude that the evidence simply is not strong enough.
And sometimes that second result is the more valuable one.
For the original research background, see Arditi et al., Refusal in Language Models Is Mediated by a Single Direction. arXiv The community implementation that helped popularize the term abliteration is also a useful reference for the basic orthogonalization idea. Hugging Face
Enjoyed this guide? Share your thoughts below and tell us how you leverage LLM abliteration in your projects!
## use Below CSS
