Ai2’s 7B moderator evaluates prompt harmfulness, response harmfulness and whether an answer is a refusal.
WildGuard’s separation of harmfulness and refusal is useful for understanding what a safety evaluation actually counts. An assistant can refuse a harmless request, answer a harmful request, or offer a helpful safe response. Those outcomes should not be collapsed into one refusal percentage.
The Ai2 card describes prompt harmfulness, response harmfulness and response-refusal detection. It identifies a 7B model based on Mistral 7B v0.3, English-language scope and an Apache 2.0 license.
Source: Allen Institute for AI · WildGuard model cardOur interpretation: a refusal label tells you how an answer responded, not whether declining was appropriate. Pair it with the prompt’s context and an assessment of the answer’s usefulness.
Include harmful requests, clearly legitimate requests and safe redirections. Inspect individual disagreements between the model and reviewers; a single average can hide a repeated failure on a particular task.
A team compares two assistant prompts. One produces more refusals. Use separate harmfulness and refusal judgments to check whether that increase blocked dangerous assistance or merely obstructed legitimate educational questions.
The English model card does not substantiate broad multilingual coverage. Results on its published evaluation data do not establish suitability for your policy or deployment.
How to read an AI safety evaluationThe sources behind this page, with a reason to open each one. Practical examples and recommendations are our editorial interpretation.
Describes prompt harmfulness, response harmfulness and refusal detection as separate tasks.
Pairs safe prompts with unsafe contrasts to investigate unnecessary refusals. Historical model results are not current rankings.
Sources reviewed 13 September 2026. Product documentation can change. How we use evidence