Why Are LLM Backdoor Defenses Fragmented? A Feature-Level Explanation with Sparse Autoencoders
Using sparse autoencoders to analyze LLM backdoors at the feature level, researchers explain why existing defenses remain fragmented and ineffective against both dirty-label and clean-label poisoning. The mechanistic analysis identifies where defense gaps originate in the model's learned representations.