Life sciences · Preprint
arXiv · September 9, 2026
Early or partial results. Treat as a signal, not a conclusion.
This is a preprint feasibility study demonstrating that directional ablation, a known white-box attack on dense language models, can be adapted to work on a frontier 320-billion-parameter mixture-of-experts model. The attack achieves 41–89 percentage-point reductions in refusal across seven harmful benchmarks, but the effect is fragmented across multiple weight modules and does not fully eliminate safety alignment; a category-concentrated residue persists even after aggressive editing.
White-box attack feasibility study on a single model. GLM-5.3-Flash (320B parameters, 288 routed experts, four-wide hyper-connection residual, block-FP8 quantization); no human subjects or held-out test population described. Intervention: Directional ablation: projection of a single refusal direction out of model weight tensors writing to the residual stream. Compared with: Ablation of a random direction orthogonal to the refusal direction (control); conventional module-name-matching recipe from prior dense-model work.
Directional ablation attack transfers to GLM-5.3-Flash (320B MoE) but does not follow the original recipe; 74% of refusal reduction emerges only from joint intervention across three module types Editing attention, dense, and routed-expert writers alone removes 0.039, 0.016, and 0.148 of refusal respectively; editing all three together removes 0.776 Conventional module-name matching accounts for only 0.066 of the 0.776 effect, explaining failure in prior MoE attempts
Attack achieves 41–89 percentage-point reductions across seven harmful benchmarks with no detected capability loss
The source did not state who this applies to in practice.
First-ever demonstration of a directional ablation attack on frontier-scale MoE models; uncontrolled, single-system study with no comparator group, establishing feasibility rather than establishing safety or efficacy claims.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Directional ablation removes an aligned language model's ability to refuse by projecting a single "refusal direction" out of the weights that write the residual stream. It needs no gradient-based training and no optimization, only a few hundred contrastive prompts, which makes it the canonical white-box attack on open-weight alignment. However, it has been established only on dense models up to roughly 70B parameters. We study whether it survives the shift to frontier mixture-of-experts (MoE) models whose residual streams are no longer a single tensor and whose weights ship quantized. We apply it to GLM-5.3-Flash (320B parameters, 288 routed experts, a four-wide hyper-connection residual, block-FP8). The attack survives the architecture, but what it reaches is no longer where a reader of the original recipe would look for it. Editing the attention, dense and routed-expert writers on their own removes 0.039, 0.016 and 0.148 of refusal respectively; editing all three together removes 0.776. As a result, 74% of the effect exists only under the joint intervention. The part the conventional recipe reaches by module-name matching accounts for 0.066 of that 0.776, which is why it fails silently on an MoE. The effect does not follow from removing just any direction: ablating a random direction orthogonal to it leaves refusal unchanged. A category-concentrated residue survives every edit we tried: subspaces fitted on violence, sexual content and hate leave measurable refusal at every rank from 1 to 12. We report the method, the 41-89 percentage-point reductions it achieves across seven harmful benchmarks with no detected change in capability, and the boundary where it stops.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.