Patrick Chao, Edoardo Debenedetti, Alexander Robey, Maksym Andriushchenko, Francesco Croce, Vikash Sehwag, Edgar Dobriban, Nicolas Flammarion et al.
An open-source benchmark to standardize and ensure reproducibility in evaluating LLM jailbreak attacks, providing state-of-the-art attack prompts, a dataset, an evaluation framework, and a leaderboard.
Existing jailbreak evaluations suffer from lack of standard practices, inconsistent cost and success rate measurements, and poor reproducibility due to closed-source code and reliance on evolving APIs.
(1) Maintain a repository of state-of-the-art jailbreak prompts; (2) Build a dataset of 100 behaviors aligned with OpenAI's usage policies; (3) Provide a standardized evaluation framework including threat model, system prompts, chat templates, and scoring functions; (4) Operate a leaderboard tracking attack and defense performance.
Enables fair and reproducible comparison of jailbreak attacks and defenses, advancing LLM safety research. Provides an evolving benchmark that the community can continuously contribute to.