A boundary is easier to understand when it preserves a path to legitimate help.
The user sees a blocked task
A model may detect a risky request. The user may see a legitimate question that went unanswered. Clear explanations can help both interpretations meet without assuming bad intent.
False positives have a cost
Blocking benign content can interrupt learning, accessibility, creative work, or professional research. Evaluate these failures alongside missed harmful requests rather than optimizing only for the number of blocks.
Leave a useful next step
An assistant can explain the boundary, ask a clarifying question, suggest a safe alternative, or offer an appropriate review path. The best choice depends on the context and the consequences of an error.
An appropriate boundary needs context
XSTest studies legitimate requests that models may refuse because they resemble unsafe ones. Its safe/unsafe contrasts help distinguish recognizing a risky topic from judging what assistance was actually requested. A sensitive subject can still have educational, protective or analytical uses.
Source: Röttger and colleagues · XSTest: identifying exaggerated safety behavioursMeasure the path after a refusal
Our view is that evaluation should include what the user can do next. Can they clarify an ambiguous request, receive a useful partial answer, or reach a qualified reviewer? A response can maintain a justified boundary and still explain the relevant limitation. Repeatedly asking the same question in frustration is not evidence that the design worked well.
Refusal rate is not a safety score
WildGuard treats refusal detection and harmfulness as separate labeling tasks. That distinction is valuable even if you never deploy it: the fact that an assistant declined says nothing by itself about whether declining was warranted. Compare errors on harmful requests with errors on legitimate ones.
Source: Allen Institute for AI · WildGuard model cardA situation to think through
A teacher asks for a discussion of propaganda techniques. A poor refusal treats the topic alone as disallowed. A more useful response can analyze historical examples without helping the user design a targeted deception campaign. The boundary depends on the requested assistance.
Questions to take with you
- Does the explanation identify the relevant boundary?
- Is there an appropriate alternative that still addresses the goal?
- Are false refusals reviewed alongside missed harmful requests?
For more reading
The sources behind this page, with a reason to open each one. Practical examples and recommendations are our editorial interpretation.
- XSTest: identifying exaggerated safety behaviours
Pairs safe prompts with unsafe contrasts to investigate unnecessary refusals. Historical model results are not current rankings.
- WildGuard model card
Describes prompt harmfulness, response harmfulness and refusal detection as separate tasks.
- OpenAI Model Spec · 18 December 2025
The provider’s intended behavior and instruction hierarchy; a policy is not proof of consistent behavior.
Sources reviewed 13 September 2026. Product documentation can change. How we use evidence