Naveen George, Naoki Murata, Yuhta Takida, Konda Reddy Mopuri, Yuki Mitsufuji
This paper proposes TILDE, a distributional alignment-based method for safely erasing unwanted concepts from text-to-image diffusion models while maintaining the quality and diversity of benign generations.
Concept unlearning is a practical requirement for deploying models under privacy, copyright, and safety regulations, necessitating the removal of specific concepts post-training. Existing methods effectively remove target concepts but often fail to explicitly preserve the 'retention' properties—quality, diversity, and semantic coverage—of benign generations, leading to performance degradation.
TILDE formulates concept unlearning as a distributional alignment problem. The target is the minimum-deviation conditional distribution from the pretrained model under a forgetting constraint. This energy-tilted, anchor-free target suppresses concept-expressing images while preserving benign relative mass for each prompt. It is instantiated with residual ∇-GFlowNet training, which learns the score correction induced by the forget energy relative to the pretrained diffusion model.
Across objects, artistic styles, and characters, TILDE achieves strong forgetting while improving retention and distributional fidelity over prior baselines. The work provides a new theoretical and practical framework for concept unlearning, contributing to the safe and practical deployment of generative models.