Brett Reynolds
This paper proposes a systematic benchmark and evaluation protocol called 'adversarial pragmatics' to diagnose complex linguistic ambiguities and instruction conflicts in language model safety evaluations.
Existing safety evaluation benchmarks often classify model behavior into simple pass/fail labels. This obscures whether failures stem from model capability limits, policy ambiguity, instruction conflicts, scaffold failures, or unstable evaluator judgments, especially when dealing with complex linguistic phenomena like embedded commands, scope ambiguity, and indirect speech acts.
The authors introduce 'adversarial pragmatics' with a linguistically controlled taxonomy covering instruction conflict, embedded commands, quotation, scope ambiguity, deixis, indirect speech acts, and multi-turn agent transcripts. They developed an 18-item seed benchmark with validator-enforced metadata, a 54-row local seed pilot, and an expert-evaluation protocol that distinguishes task success, policy compliance, safety risk, refusal outcome, and evaluator confidence.
The framework turns linguistic judgment methodology into a practical tool for validating safety evaluations, LLM judges, gold-set construction, prompt-injection tests, and safety documentation. Its key contribution is providing a methodological foundation to move beyond binary evaluations and precisely diagnose the root causes of failures in AI safety assessments.