Lijun Li, Bowen Dong, Ruohui Wang, Xuhao Hu, Wangmeng Zuo, Dahua Lin, Yu Qiao, Jing Shao
SALAD-Bench is a large-scale hierarchical benchmark for LLM safety evaluation, covering attack and defense methods as well.
Existing LLM safety benchmarks are limited in scale or diversity, and fail to integrate evaluation of attack and defense methods. Moreover, reliable evaluation of complex attack-enhanced queries has been difficult.
SALAD-Bench constructs a large-scale question set with a three-level taxonomy, and adds attack and defense modifications to generate questions of varying difficulty. For evaluation, it introduces an LLM-based MD-Judge to reliably evaluate QA pairs even with attacks.
SALAD-Bench unifies LLM safety evaluation, attack method evaluation, and defense method evaluation into a single benchmark, enabling researchers to comprehensively assess LLM resilience against various threats and the effectiveness of defense strategies.