TRAP: Mitigating Poisoning-based Backdoor Attacks by Treating Poison with Poison

Chun Li, Zhong Li, Minxue Pan, Xuandong Li

IEEE Transactions on Dependable and Secure Computing
TDSC 25 · Dec 2025

Abstract

The backdoor attack poses a significant threat to deep neural networks. Existing works on poison suppression defense mainly focus on differentiating between poisoned and benign samples based on various metrics and removing the backdoor using Unlearning. However, these metrics can be bypassed by certain attacks, and Unlearning often leads to sub-optimal model performance. Through examination of the model attack process, we discovered that poisoned samples always form clusters distant from benign samples in the early stages of training and when the model is fully trained, there is a unique pathway within the classifier connecting backdoor features to the target label. Leveraging these observations, in this paper, we propose a novel training method to detect poisoned samples during the early stages of training and remove the backdoor by retraining the classifier part of the model on relabeled poisoned samples. We evaluated our method against twelve attacks on four datasets, and the results showed that our method significantly outperforms existing state-of-the-art defenses. We reduced the average attack success rate to 0.07% while only decreasing the average accuracy by 0.33%. Our code is available at https://anonymous.4open.science/r/TRAP-2672.

Backdoor attacksTrustworthy machine learningSecurity
First page of TRAP: Mitigating Poisoning-based Backdoor Attacks by Treating Poison with Poison
TDSC 25 · PDF

BibTeX